Trending: On-device modelsSearch
iHeartGeek
iTECH

AWS and Nvidia add 2 million GPUs and Vera CPUs

The two companies have expanded a 16-year partnership with 2 million extra Nvidia GPUs, Nvidia Vera CPUs on AWS and closer work on the machines both call physical AI.

A long aisle of dark server racks receding into the distance inside a data centre, lit by cool blue light with a warm glow spilling from a doorway at the far end

Two million more GPUs. AWS and Nvidia have expanded a partnership that both companies describe as 16 years old, with a plan to deploy 2 million additional Nvidia GPUs across AWS's global infrastructure and to push deeper into AI factories, CPUs, networking, open models and software.

The two chief executives signed their names to it. "Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together," said Matt Garman, CEO of AWS. Jensen Huang, founder and CEO of Nvidia, called the pair "one of the great growth engines of the AI era" and said demand is running ahead of every forecast.

What is actually being built

The GPU figure stacks on top of what AWS already announced at Nvidia GTC 2026, when it said it would add more than a million Nvidia GPUs starting in 2026. The new 2 million are described as Blackwell Ultra and Rubin class hardware, and the companies say customer demand exceeded the earlier plan rather than the plan being revised down.

There is specific silicon in the detail. AWS plans to expand Blackwell capacity with Nvidia RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances, which AWS says deliver 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 generation, and it claims to be the first major cloud provider to offer instances accelerated by that part. AWS and Nvidia are also working to bring Vera CPU-based infrastructure to AWS, aimed at agentic AI workloads that need serious CPU compute alongside accelerators.

The memory story is tucked into the same announcement. The two companies are extending Nvidia NVLink Fusion support, announced at re:Invent 2025 for next-generation Trainium chips, so that it works with Nvidia's new NVHBM custom high-bandwidth memory, in partnership with memory suppliers. Given how hard high-bandwidth memory has been to buy this year, wiring an AWS-designed chip into Nvidia's interconnect and memory standard is a supply argument as much as an engineering one.

Why robots keep turning up in GPU announcements

The framing this time is agentic and physical AI, which is the industry's shorthand for software that takes actions and models that drive machines rather than answer questions. Translation: the workloads AWS and Nvidia expect to grow into are ones that keep a GPU busy in a loop, whether that is an agent chewing through a data pipeline or a robot deciding where to put its hand. Those are exactly the jobs that cannot be batched cheaply and queued politely, which is why the pitch is more capacity rather than a cleverer chip.

Our opinion

Read this as a capacity forecast wearing a press release, which is what these announcements always are. There is no contract value, no delivery schedule and no regional split, so nobody outside the two companies can check whether 2 million is a plan or an ambition, and "demand exceeded our expectations" is doing the work that a signed order book would normally do. The genuinely interesting line is the memory one. Nvidia's NVHBM memory standard turning up inside AWS's own Trainium silicon is a concession from both sides: Nvidia accepts that the big clouds will keep building their own chips, and AWS accepts that the interconnect and memory hierarchy around those chips is going to be Nvidia's. That is a more durable arrangement than a GPU count, and it is the part of this announcement that will still matter in three years when the next 2 million are old news.