Trending: On-device modelsSearch
iHeartGeek
iTECH

CoreWeave puts NVIDIA's Vera Rubin into production

CoreWeave has made NVIDIA's Vera Rubin NVL72 available on its cloud with Spectrum-X networking, and Cognition is already running its Devin agent on the new systems.

NVIDIA's Vera Rubin NVL72 rack system, showing the lit compute trays inside a single tall cabinet

CoreWeave has moved NVIDIA's next generation of data centre hardware into production, announcing availability of Vera Rubin NVL72 systems alongside 102.4T Spectrum-X Ethernet networking at its Fully Connected event in San Francisco.

Cognition gets the first production slot

Cognition, the applied AI lab behind the Devin software engineer, is the first customer running production workloads on the new systems. Benchmarking against a GB200 NVL72 baseline on a real software engineering workload drawn from FrontierCode, the company reported up to a 4.8x increase in total token throughput for SWE-2 inference workloads. Cognition had scaled to thousands of GPUs on CoreWeave within nine months.

"Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship," said Silas Alberti, a member of Cognition's founding team. NVIDIA's Ian Buck framed the pitch as durability rather than raw numbers, noting that CoreWeave's V100 GPUs are still running customer workloads nearly a decade after Volta launched.

Vera CPU and a platform for agents

CoreWeave will also offer NVIDIA Vera, the first CPU the company has built specifically for AI agents. A single rack carries 128 CPUs and 11,264 cores, which CoreWeave says is enough for more than 11,000 isolated environments at one core each. In its testing it recorded more than three times faster agent sandbox start-up times and a 1.7x gain across all passing tasks on Terminal-Bench.

Alongside the hardware, CoreWeave launched Forge, a connected environment for training, evaluating and improving models that bundles Weights & Biases, post-training tooling from OpenPipe and the open source marimo notebook project. Canva, Capital One and MasterClass are among the first companies building on it. CoreWeave also took its sandbox service and the ARIA assistant to general availability and introduced Agent Lens for production agent observability, while NVIDIA's Dynamo framework powers its managed inference service and an RL Rollouts preview.

Our opinion

The number that matters is not the 4.8x token throughput, it is the nine months Cognition needed to reach thousands of GPUs. Agents in production behave like a demand curve rather than a load estimate, and infrastructure that cannot be added in weeks is infrastructure that quietly becomes the reason a product ships late.

Vera's rack density also reframes what agent infrastructure is for. Sandbox count, not inference throughput, is the constraint the moment a lab wants thousands of parallel environments for reinforcement learning, and building a CPU around that specific bottleneck is a more honest signal of where the industry thinks agent training is heading than any benchmark chart.