CoreWeave Opens Nvidia Vera Rubin NVL72 to Customers, With Devin's Maker Reporting 4.8x Token Throughput. What the Number Does and Doesn't Mean
Cognition measured the gain on its own SWE-2 inference workload against a GB200 NVL72 baseline. Pricing, regions, and results for other workloads were not given.
TL;DR
On September 30, 2026, CoreWeave said Nvidia's Vera Rubin NVL72 systems are available on CoreWeave Cloud, with Cognition, the maker of the Devin coding agent, as the first customer running production workloads. Cognition's engineers measured 4.8 times the total token throughput on their SWE-2 inference workload and 3.8 times the output token throughput on reinforcement-learning workloads, both against a GB200 NVL72 baseline. The figures come from one customer's own agentic coding workload, and the announcement gives no pricing or regional availability.
Coding agents like Devin make many model calls per task, so the cost and speed of inference sets how many sessions a company can run at once. CoreWeave's announcement is about that bottleneck.
What was announced
- Nvidia Vera Rubin NVL72 is now available on CoreWeave Cloud, announced on September 30, 2026 alongside CoreWeave's Fully Connected 2026 conference.
- Cognition is the first customer to run production agentic workloads on it.
- Cognition's engineers ran the benchmark themselves on CoreWeave Cloud, comparing against their existing GB200 NVL72 cluster.
- CoreWeave's Chen Goldberg said the investment shows up in "the ability to get production workloads running within days."
The numbers
| Workload | Result vs GB200 NVL72 | Measured by |
|---|---|---|
| SWE-2 inference (total token throughput) | 4.8x | Cognition |
| Reinforcement learning (output token throughput) | 3.8x | Cognition |
What the number does not tell you
- It is one customer's agentic coding workload. CoreWeave itself notes gains may vary for other uses.
- It is throughput, not price. The announcement gives no cost per token, so cost per Devin session cannot be worked out from it.
- No regions or capacity numbers were given, and availability is described only as available on CoreWeave Cloud.
- It is a comparison with the previous Nvidia generation, not with other chips or clouds.
What this means for people choosing coding agents
Nothing changes in the Devin product you use today. The practical signal is that agent vendors are tuning for throughput per GPU, which is what eventually shows up as more concurrent sessions or lower prices. Judge a coding agent on the results you get on your own repository, and watch for any pricing change from Cognition rather than assuming this one.
AI Industry Reporter
Priya covers model releases, industry announcements, and the gap between what labs claim and what independent evaluators actually find. She reads the primary source - the paper, the system card, the benchmark org's own statement - before writing a word.
More on AI Coding Assistants
DeepSeek V4.1-Flash Replaces V4-Pro at $0.60 per Million Output Tokens: The Price Table, the Benchmarks DeepSeek Reports, and the Gaps It Admits
Priya Nair · 5 min
OpenAI and Synopsys Announce GPT-Synopsys, a Model That Runs Chip-Design Software. What's Confirmed, and What Isn't Priced or Dated Yet
Priya Nair · 5 min
Gemini 4 Argon Is Announced, but Only Trusted Cyber Defenders Can Use It. Here's the Price, the Benchmarks Google Reports, and What's Still Missing
Priya Nair · 5 min