Skip to content
TECH CEO Daily
Big TechAnalysis

CoreWeave Vera Rubin NVL72 enters production with Cognition, reporting up to 4.8x gains

CoreWeave says Cognition is the first customer anywhere running production work on Nvidia's Vera Rubin racks, and it plans to rent Nvidia's new Vera CPUs for AI agents.

By · Editor

· 3 min read · Fact-checked

The 60-second brief

  • 1CoreWeave said on September 30 that Cognition is the first customer running production workloads on Nvidia Vera Rubin NVL72.
  • 2Cognition reported up to 4.8x the token throughput per GPU of a GB200 NVL72 baseline on its inference workload.
  • 3CoreWeave will also offer Nvidia's Vera CPU, built for AI agents, with 11,264 cores per rack; no date given.

The news

CoreWeave (CRWV) said on September 30, 2026, that its Vera Rubin NVL72 systems, Nvidia's next-generation rack of 72 GPUs, are running production work, with AI coding company Cognition as the first customer anywhere to do so. The announcement came at CoreWeave's Fully Connected 2026 conference in San Francisco.

According to CoreWeave, it stood up the Vera Rubin cluster in early September and had Cognition's workloads running within days of the racks being handed over. Cognition, which makes the Devin coding agent, runs training, reinforcement learning and production inference on CoreWeave, and the company said Cognition scaled from bridge capacity to thousands of GPUs in less than nine months.

Cognition's engineers reported up to 4.8 times the total token throughput per GPU of a GB200 NVL72 baseline on inference for its SWE-2 model, and 3.8 times the output token throughput per GPU on reinforcement learning, CoreWeave said. StorageReview, which covered the launch, noted that no model sizes, context lengths or batch settings were disclosed, so the figures cannot be compared directly with other benchmarks. StorageReview also reported the service is in limited availability with gated allocation.

In a second announcement the same day, CoreWeave said it will offer Nvidia's Vera CPU, which Nvidia pitches as the first processor designed for AI agents. CoreWeave described racks of 128 CPUs and 11,264 cores that can support more than 11,000 isolated agent environments, and claimed three times faster sandbox startup than x86 alternatives. StorageReview reported the CPU service is listed as coming soon, without a date.

Rubin demand is also building outside the large US clouds. AM Intelligence, the AI infrastructure arm of India's AM Group, said in August that it had placed a firm order for 9,000 Nvidia Rubin GPUs in Vera Rubin NVL72 systems for a 30-megawatt site in Hyderabad, with delivery expected in the first quarter of 2027, according to Entrepreneur India.

The numbers

Inference throughput per GPU vs GB200 NVL72 (Cognition, SWE-2)
Up to 4.8x
RL output token throughput per GPU
3.8x
Cores per Vera CPU rack
11,264
Isolated agent environments per rack
11,000+
AM Intelligence Rubin GPU order (Hyderabad)
9,000

Why CEOs should care

For technology buyers renting AI capacity, the message is that a new hardware generation is already in customers' hands. If a vendor's own customer reports several times the throughput per GPU, the price per token on older GB200 capacity could come under pressure. Buyers signing multi-year GPU commitments should ask for upgrade or swap rights to newer racks, and for pricing tied to throughput rather than GPU-hours alone.

CFOs should treat vendor benchmark multiples with care. Cognition's 4.8x figure came from its own workload, and the settings were not disclosed. Before budgeting savings, ask providers to run your models on the new hardware and to put measured results, not headline multiples, into the contract.

The Vera CPU launch points to a cost line many boards have not tracked: the general-purpose computing that AI agents use to run code, browse and call tools. CIOs running agent pilots should measure how much of their spend sits on CPUs and sandboxes, not only GPUs, and ask whether isolated agent environments meet their security rules.

The bigger picture

Neoclouds such as CoreWeave compete with Amazon, Microsoft and Google by getting each Nvidia generation into production quickly. Being first with Rubin gives CoreWeave a marketing edge, while orders like AM Intelligence's show national and regional providers are also lining up for the same supply. Early access tends to go to large, fast-scaling customers such as Cognition, which means smaller buyers may wait longer for allocations.

The two announcements also show where AI spending is heading. GPUs still dominate budgets, but CoreWeave and Nvidia are now selling processors and isolated environments built for agents, which run many small tasks around each model call. Signal65 president Ryan Shrout, quoted by CoreWeave, argued that the CPU share of agent work will keep growing. If that holds, the next round of capacity deals will cover full racks of mixed hardware rather than GPU counts alone.

What’s next

Watch for when CoreWeave moves Vera Rubin from limited availability to general availability, a launch date for its Vera CPU service, and whether other customers publish measured results. AM Intelligence's Hyderabad deliveries are scheduled for the first quarter of 2027.

What “Fact-checked” means

Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.

What we checked
Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
How
A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
Who
The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, . A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
If something is wrong
“Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error

How we fact-check →

Companies in this story

CoreWeaveNvidiaCognitionVera RubinAI infrastructure

Earlier coverage of NVIDIA

All NVIDIA coverage →

Written by

Editor · Technology & Business Writer

Hussein is a writer and business technology enthusiast focused on the intersection of technology, entrepreneurship, finance, artificial intelligence, and digital innovation.

CoversAICybersecurityBig TechSaaSStartupsFintech

About this story. Researched from primary sources whenever they are available and fact-checked before publication.

Published by Tech CEO Daily, an independent publication. Masthead · Editorial standards

Follow Tech CEO Daily on Facebook for the day’s top stories in your feed.

Free newsletters

The technology briefing for people running businesses.

Daily, weekly, bi-weekly or monthly. You choose.

How often

The Daily Brief · Monday to Saturday, 7 a.m. ET

Free forever. One click to unsubscribe. We never sell your email.