Nvidia backs AI power efficiency push with new deals, ETDatacenters


Nvidia has outlined a series of collaborations and performance results focused on enhancing AI infrastructure efficiency, particularly how AI systems consume electricity amid increasing demand for agentic AI workloads. Speaking at the AI Infra Summit in Santa Clara, Ian Buck, Vice President of Hyperscale and High-Performance Computing at Nvidia, stated that the industry’s focus is shifting from peak chip performance to output measured against power consumption. This transition is making token throughput per megawatt a critical metric for AI deployments, especially as larger models and agent-based applications strain data centre power budgets.

Power Focus and Benchmarks


Among the announcements, Emerald AI, Nvidia, and Silicon Valley Power collaborated on a commercial AI factory flexible-load programme. This initiative is designed to reduce electricity demand when required by the grid while ensuring priority AI tasks continue to run. The system successfully responded to hundreds of utility demand signals by throttling lower-priority AI jobs during periods of grid stress and restoring normal operations later. Nvidia suggested this approach could enable AI facilities to function as controllable loads, potentially allowing utilities to support more computing sites without immediate expansion of electricity infrastructure.Lambda, an AI cloud provider, also released test results for Nvidia’s DSX MaxLPS software on Blackwell servers. The figures indicated Lambda ran 19 nodes within a power budget typically allocated for 16 full-power nodes. This configuration increased cluster-wide token throughput by 24 per cent, from approximately 4 million to 5 million tokens per second, while improving performance per watt by 23 per cent. For operators seeking to expand output without securing additional power, these results underscore the growing importance of electricity as a constraint in the AI infrastructure market, alongside chip supply.Nvidia also highlighted its Vera Rubin NVL72 as a central system for large-scale AI deployments, claiming DSX MaxLPS can provide up to 40 per cent more GPU capacity within the same megawatt budget in suitable environments and up to 35 per cent higher token throughput without requiring new power lines. Additionally, Nvidia linked Vera Rubin to Groq 3 LPX for inference workloads, stating the combined setup can deliver up to 35 times higher token throughput per megawatt than GB200 NVL72 for very large models with long context windows. Benchmark claims on SemiAnalysis AgentX indicated Vera Rubin NVL72 delivered up to 30 times higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model.

Platform Links and Reliability


Nvidia also disclosed several platform partnerships. Amazon’s Annapurna Labs is collaborating with Nvidia on NVHBM custom high-bandwidth memory technology, while d-Matrix is integrating with NVLink Fusion to combine Nvidia Vera CPUs with d-Matrix Raptor XPUs for low-latency inference. Pinterest is utilising the Nvidia Blackwell platform and Nvidia Dynamo inference software for conversational AI in visual discovery.As AI factories scale to hundreds of thousands of GPUs, Nvidia emphasised that reliability is increasingly critical for economic performance. The company introduced NVLink 6, a system designed to detect and isolate faults before they affect applications. This design incorporates error correction, retry mechanisms, dynamic routing, and link rebalancing, intended to prevent local failures from propagating across large AI clusters.“The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt,” Buck said. India’s emerging AI data centers and hyperscale deployments will need to evaluate current and future AI infrastructure plans for power consumption and efficiency metrics. Indian data center operators should investigate flexible-load programs and software solutions for managing AI workloads in response to grid signals, and assess the implications of new high-density compute architectures and associated power and cooling requirements for data center design and upgrades.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *