# AWS and Nvidia Expand Historic AI Infrastructure Partnership With Multi-Billion-Dollar GPU Commitments
Amazon Web Services has announced a sweeping expansion of its artificial intelligence hardware roadmap, revealing plans to deploy an additional two million Nvidia GPUs across its worldwide data centre footprint over the coming years. The move underscores the enormous scale of investment flowing into AI-ready cloud infrastructure as demand for generative and agentic AI compute continues to outpace supply.
## A Massive Expansion of Nvidia-Powered Capacity
The newly disclosed commitment adds to an earlier agreement announced at Nvidia’s annual GTC conference earlier in the year. That initial deal covered more than one million GPUs set to begin arriving in 2026. Together, the two agreements represent a planned deployment of more than three million Nvidia GPUs across AWS infrastructure by the end of the deployment window.
The latest batch of hardware includes Nvidia’s newest GPU architectures: Blackwell Ultra, Rubin, and Rubin Ultra. These chips represent successive generations of Nvidia’s AI accelerator lineup, each offering significant performance improvements for training and inference workloads. The agreement extends beyond GPUs to encompass Nvidia CPUs, networking hardware, and advanced interconnect technologies designed to tie everything together at scale.
A notable component of the deal is a dedicated allocation of 100,000 Nvidia GPUs to be deployed on secure AWS infrastructure specifically for United States federal government and national-security workloads. The companies described these systems as part of a broader AI infrastructure initiative being developed on behalf of the US government, signalling the growing role of cloud AI platforms in national defence and intelligence operations.
## AWS Capacity Already Booked Years Into the Future
The announcement arrives at a time when AWS is already facing a surge in advance customer bookings that extends deep into the coming years. Amazon CEO Andy Jassy noted that the vast majority of AWS compute capacity slated for 2027 has already been reserved by customers, and reservation activity is also reaching into 2028.
This strong demand is reflected in AWS’s financial metrics. The company’s contracted backlog at the end of the second quarter of 2026 stood at $496 billion, a sharp increase from $364 billion just three months prior. During the same quarter, AWS generated $42.2 billion in revenue, representing year-over-year growth of 37 percent.
Despite the impressive scale of planned additions, Jassy cautioned that supply would continue to lag demand. “Even at that amount, we will still not have enough capacity to meet all of the demand we have in 2026,” he said, highlighting the persistent gap between infrastructure buildout and the insatiable appetite of AI workloads.
## Soaring Capital Expenditure Reflects Infrastructure Scale
The cost of building this infrastructure is staggering. Amazon’s cash capital expenditure during the second quarter of 2026 reached $53.1 billion, nearly double the $31.4 billion spent in the same quarter of the prior year. For the first half of 2026 alone, cash capital expenditure totalled $96.3 billion, compared to $55.6 billion in the same period a year earlier.
Amazon attributed the spending primarily to technology infrastructure investments that support AWS growth, along with additional capacity for its logistics and fulfilment network. The company subsequently raised its full-year 2026 capital expenditure forecast from approximately $200 billion to roughly $220 billion, citing higher memory costs as a key driver of the increase.
The broader picture across the technology sector is equally striking. Combined 2026 capital expenditure estimates for Amazon, Microsoft, Alphabet, Meta, and Oracle jumped from approximately $485 billion in January to around $730 billion by July, according to analysis of financial data. These figures encompass all infrastructure spending, as the companies do not consistently separate AI-specific investments from general technology expenditure.
AWS’s infrastructure investments involve considerable lead times. The company typically commits capital for land, power, buildings, chips, servers, and networking equipment between six and 24 months before the resulting capacity begins generating billable revenue. Data centres themselves have useful lives exceeding 30 years, whereas the chips, servers, and networking equipment inside them tend to have operational lifespans of just five to six years. Jassy noted that Amazon can begin spending on a new data centre roughly two years before the facility opens its doors.
## Trainium: AWS’s Own AI Chip Portfolio Grows in Tandem
While deepening its Nvidia partnership, AWS is simultaneously expanding its proprietary AI processor portfolio. The company’s Trainium3 processor began shipping at the start of 2026, with Trainium4 expected to enter deliveries in 2027.
Trainium3 quickly became the centre of strong customer interest, with AWS reporting in May that the chip was nearly fully subscribed. Much of the planned Trainium4 capacity has also already been reserved by customers. The company disclosed more than $225 billion in revenue commitments tied to Trainium, signalling deep market confidence in its in-house silicon strategy.
Designed for both AI training and inference workloads, Trainium gives AWS an alternative to relying on third-party accelerators. Amazon developed the processor through Annapurna Labs, the chip design firm it acquired in 2015. At scale, the company expects Trainium to deliver savings of tens of billions of dollars in annual capital expenditure and to provide several hundred basis points of operating-margin advantage compared with deploying other chips for inference tasks.
The relationship between Trainium and Nvidia hardware is also deepening. Trainium4 is being engineered to support Nvidia’s NVLink Fusion interconnect technology, which enables high-bandwidth connections between processors and accelerators within AI systems. The collaboration between the two companies extends to multiple layers of the technology stack, with the earlier agreement including Nvidia ConnectX and Spectrum-X networking equipment, and the newer agreement covering work to bring Vera CPU-based infrastructure to AWS and expand NVLink Fusion support through Nvidia’s custom high-bandwidth memory technology.
Nvidia and Amazon’s Annapurna Labs jointly stated that the memory and interconnect efforts are intended for future Trainium-based infrastructure, enabling Trainium accelerators and Nvidia GPUs to operate within a shared rack-scale architecture. This hybrid approach allows AWS to maximise the strengths of both its proprietary silicon and Nvidia’s industry-leading GPUs within the same data centre environment.
## Nvidia’s Own Growth and Constraints
Nvidia’s role in the AI infrastructure boom is reflected in its own financials. The company reported $89 billion in data centre revenue for its fiscal second quarter, more than double the year-ago figure. Yet Nvidia itself acknowledged that it is operating at the limits of its supply chain.
“We are supply-constrained,” Nvidia CFO Colette Kress stated during the company’s earnings call. The company also warned that higher memory and component costs would continue to weigh on its profit margins, even as demand remains exceptionally strong.
Nvidia noted that specialist AI cloud providers such as CoreWeave and Nebius are expected to finish 2026 with more than eight gigawatts of Nvidia GPU capacity, up from three gigawatts at the end of 2025. The Vera Rubin platform has also begun shipping to customers, with AWS, Google Cloud, Microsoft, and Oracle Cloud Infrastructure among the first cloud providers named as expected deployers of Vera Rubin-based instances.
## The Energy Challenge Behind AI Scaling
All of this GPU capacity requires a corresponding expansion of the physical infrastructure that supports it — particularly in terms of power and cooling. The International Energy Agency projects that global data centre electricity consumption will rise from approximately 415 TWh in 2024 to roughly 945 TWh by 2030, with accelerated servers accounting for nearly half of that projected increase.
Electricity consumption driven by accelerated servers, largely fuelled by AI adoption, is expected to grow by roughly 30 percent annually under the IEA’s base case scenario. Cooling systems and other supporting infrastructure are projected to account for an additional 20 percent of the increase in data centre electricity use over the same period. The IEA estimates that data centres can become operational within two to three years, but the energy infrastructure needed to power them generally requires substantially longer planning and construction timelines.
Amazon has also been navigating power constraints more directly, with a separate article noting that the company is shifting AWS workloads in response to tightening power availability in certain regions. This reality adds another layer of complexity to the already immense challenge of deploying millions of GPUs in a timely fashion.
—
## Frequently Asked Questions
**What does AWS’s new Nvidia GPU commitment cover?**
The commitment includes the deployment of two million Nvidia GPUs across AWS’s global infrastructure in 2027 and 2028. It encompasses Blackwell Ultra, Rubin, and Rubin Ultra GPUs, as well as Nvidia CPUs, networking equipment, and interconnect technology. A separate allocation of 100,000 GPUs is designated for secure US federal and national-security workloads.
**How many total Nvidia GPUs is AWS planning to deploy?**
Across both the earlier commitment (more than one million GPUs starting in 2026) and the newly announced agreement (two million GPUs in 2027 and 2028), AWS has more than three million Nvidia GPUs planned for deployment.
**What is Trainium and why does it matter?**
Trainium is AWS’s proprietary AI processor designed for training and inference workloads. It is developed through Annapurna Labs, which Amazon acquired in 2015. Trainium provides AWS with an alternative to third-party accelerators, and at scale it is expected to save tens of billions of dollars in annual capital expenditure while delivering significant operating-margin advantages.
**How does Trainium work alongside Nvidia hardware?**
AWS is building a hybrid infrastructure strategy. Trainium4 is being designed to support Nvidia’s NVLink Fusion interconnect, enabling Trainium accelerators and Nvidia GPUs to operate within a common rack-scale architecture. This allows AWS to leverage the strengths of both its own silicon and Nvidia’s GPUs.
**Why is AWS raising its capital expenditure so dramatically?**
Amazon raised its 2026 capex forecast from roughly $200 billion to about $220 billion, primarily driven by higher memory costs and the enormous scale of infrastructure investment required to meet AI demand. The company’s capex for the first half of 2026 alone reached $96.3 billion, nearly double the same period a year earlier.
**What is driving the unprecedented demand for AI compute?**
Customer reservations for AWS compute capacity now extend into 2028, and the contracted backlog grew from $364 billion to $496 billion in a single quarter. The rapid growth of generative AI, agentic AI, and enterprise AI adoption is fuelling demand that currently outpaces the supply of available compute capacity.
**What are the energy implications of this AI infrastructure buildout?**
The IEA projects global data centre electricity consumption will roughly double from 415 TWh in 2024 to 945 TWh by 2030. Accelerated servers account for nearly half of that increase, with cooling and other infrastructure adding another 20 percent. Energy infrastructure planning timelines are longer than the data centre construction timelines, creating a potential bottleneck.
—
## Conclusion
AWS’s expanding partnership with Nvidia represents one of the largest coordinated infrastructure investments in the history of cloud computing. With more than three million GPUs committed across two agreements, billions of dollars in additional capital expenditure, and a parallel investment in its own Trainium silicon, AWS is positioning itself at the centre of the AI compute revolution.
At the same time, the scale of these commitments exposes the significant challenges involved — from supply chain constraints at Nvidia to power availability and the long lead times required to build data centre infrastructure. The gap between reserved capacity and projected demand shows no signs of narrowing, and companies across the technology sector are making similarly aggressive investment bets.
The coming years will reveal whether this unprecedented level of investment translates into sustained competitive advantage and the ability to meet the world’s growing appetite for artificial intelligence. What is already clear is that the era of trillion-dollar infrastructure commitments is well underway, and AWS and Nvidia are leading the charge.
Thank you for reading



