NVIDIA Vera Rubin is the company's data center platform after Blackwell, pairing the Rubin GPU with the Vera CPU, as described in NVIDIA's 5 January 2026 announcement. Production shipments began in August 2026, CFO Colette Kress said on 26 August 2026. CoreWeave announced Vera Rubin NVL72 availability on its cloud on 30 September 2026; NVIDIA's post that day described the offering as accessible to early-access customers. NVL72 is the rack's current name; NVIDIA renamed it from NVL144, as dated in the next section.
As of 1 October 2026, this site's live price tables do not list Rubin yet, so the rent-now section shows Blackwell, with H200 alongside it. A shipment announcement is useful planning evidence. It does not tell you which allocation you can book, what software it includes or whether the provider will commit to your start date.
Decode the chip, superchip and rack names
NVIDIA's blog dated 13 October 2025 now carries this editorial note: "This blog has been updated to reflect a branding change from Vera Rubin NVL144 to Vera Rubin NVL72." That is the page's publication date, not a dated notice of when the edit happened. SemiAnalysis explained on 25 February 2026 that the old name counted 144 compute dies in 72 two-die GPU packages. The new name counts packages; SemiAnalysis dates the reversal to late December 2025.
All Vera Rubin NVL72 references below use that package-count convention. Do not treat the older name as evidence of twice as many GPUs. Keep the product name and its date together when reading an older rack proposal.
| Name | Meaning and dated attribution |
|---|---|
| Vera | The CPU paired with Rubin; NVIDIA, 5 January 2026. See the Vera CPU guide. |
| Rubin | The GPU, with two compute dies unified in one package by NV-HBI; NVIDIA architecture article, 21 July 2026. |
| VR200 | Reported rack shorthand in Tom's Hardware's VR200 NVL72 coverage, 21 May 2026; not confirmation of a standalone GPU SKU. |
| Vera Rubin Superchip | Two Rubin GPUs and one Vera CPU; NVIDIA datasheet, as of 1 October 2026. |
| Vera Rubin NVL72, formerly NVL144 | The rack with 72 Rubin GPUs and 36 Vera CPUs; NVIDIA DGX page, as of 1 October 2026. SemiAnalysis, 25 February 2026, says the 72 counts GPU packages; renaming explained above. |
| HGX Rubin NVL8 | Eight Rubin SXM GPUs on a baseboard; NVIDIA HGX page, as of 1 October 2026, distinguishes x86 and single-socket Vera versions. |
| DGX systems on Rubin | NVIDIA's complete DGX Vera Rubin NVL72 rack and DGX Rubin NVL8 system; NVIDIA product pages, as of 1 October 2026. |
| Rubin CPX | Context-processing GPU announced by NVIDIA on 9 September 2025; Ian Buck later said NVIDIA had pulled it to prioritize LPU decode in 2026, in the transcript published 23 March 2026. See CPX and Groq LPX. |
| Rubin Ultra | A later generation, targeted for the second half of 2027 in NVIDIA's 25 March 2025 recap; a historical target, not a shipment statement. See the NVIDIA GPU roadmap. |
For purchasing documents, write the full configuration. "NVIDIA VR200" alone leaves too much unspecified. The Superchip, an HGX baseboard and a complete DGX rack are different products. Ask for the GPU package count, CPU type, memory specification and network allocation in the same quote.
NVIDIA's HGX page, as of 1 October 2026, distinguishes HGX Vera Rubin NVL8 with a single-socket Vera CPU from an x86 variant whose host specifications belong to the OEM. NVIDIA's DGX Rubin NVL8 page, as of the same date, lists two Intel Xeon 6776P processors and marks its specifications preliminary. The Rubin name therefore does not identify the host CPU in every system. If your deployment requires a particular CPU environment, specify that requirement before comparing offers.
NVIDIA Rubin specs next to Blackwell Ultra
Both columns below are NVIDIA's own figures for one GPU package. The Rubin column comes from NVIDIA's 21 July 2026 architecture article and its Vera Rubin NVL72 product page, as of 1 October 2026. The Blackwell Ultra column comes from NVIDIA's 22 August 2025 Blackwell Ultra architecture article and the comparison table in its Rubin technical article, published 5 January 2026 and updated 16 March 2026. Compute figures are VENDOR CLAIM peak rates.
| Per GPU | Rubin | Blackwell Ultra |
|---|---|---|
| Compute dies | Two, joined in one package by NV-HBI | Two |
| Transistors | 336 billion | 208 billion |
| Memory | Up to 288 GB HBM4 | Up to 288 GB HBM3E |
| Memory bandwidth | Up to 22 TB/s peak (July article); 19.2 TB/s (product page) | 8 TB/s |
| NVFP4, dense | 35 PFLOPS, labelled training | 15 PFLOPS |
| NVFP4, sparse | 50 PFLOPS, labelled inference | 20 PFLOPS |
| FP8, dense | 17.5 PFLOPS, FP8/FP6 training | 5 PFLOPS |
| NVLink per GPU | NVLink 6: 3,600 GB/s (July article); 3 TB/s (product page) | 1,800 GB/s |
| NVLink-C2C to the CPU | 1.8 TB/s to Vera | 900 GB/s to Grace |
By our arithmetic, Rubin has about 2.3 times the dense NVFP4, 2.5 times the sparse NVFP4 and 3.5 times the dense FP8 of Blackwell Ultra. VENDOR CLAIM: NVIDIA's July article calls the 22 TB/s figure 2.8 times Blackwell Ultra's 8 TB/s. These are ratios of peak specifications, not forecasts of how fast your job runs. Memory capacity does not grow: both top out at 288 GB per GPU, so the gain is bandwidth and compute, not room for a larger model.
Keep each comparison inside one source. This site's GPU specification reference lists the B300 at 13,500 dense and 18,000 sparse FP4 TFLOPS and 4,500 dense FP8 TFLOPS, figures taken from NVIDIA's HGX page. They differ from the architecture-article figures above, so never divide a Rubin number by one of them.
NVIDIA's two Rubin bandwidth pairs remain unresolved: 22 TB/s and 3,600 GB/s in the 21 July 2026 article, 19.2 TB/s and 3 TB/s on the product page as of 1 October 2026. Ask a provider which figure its delivered system meets rather than budgeting on the larger one.
Two more figures are reports, not NVIDIA specifications. ServeTheHome reported on 5 January 2026 that Rubin's compute dies are made on TSMC's 3nm process. ANALYST ESTIMATE: SemiAnalysis described Rubin operating configurations of 2,300 W (Max-P) and 1,800 W (Max-Q) on 25 February 2026. That is not a universal NVIDIA TDP, so take any hosting power figure from the system supplier.
For compute, keep all three labels: precision, sparsity and workload. NVIDIA's January technical article calls the 50 PFLOPS inference headline Transformer Engine compute and labels the 35 PFLOPS training figure dense; the product page labels the inference figure sparse. Do not halve one to get the other. The HBM4 guide covers the memory change in more detail.
The rack combines an NVLink domain with a wider network
NVIDIA's DGX page, as of 1 October 2026, lists 72 Rubin GPUs, 36 Vera CPUs and 20.7 TB HBM4 in DGX Vera Rubin NVL72. It explicitly marks the specifications preliminary and subject to change. NVIDIA's 16 March 2026 technical article describes the 72 GPUs as one scale-up domain connected through an NVLink copper spine.
NVIDIA's datasheet, as of 1 October 2026, also advertises up to 75 TB of fast-access memory, including CPU memory. That is not an HBM-only capacity. Keep GPU memory and host memory separate when sizing a model. The full rack total is also not a promise that your rental allocation includes the whole rack.
| Rack item | Published figure and qualification |
|---|---|
| NVFP4 inference, sparse | VENDOR CLAIM: 3,600 PFLOPS, NVIDIA NVL72 page, as of 1 October 2026; 3.6 EFLOPS by our arithmetic, 3,600 / 1,000 |
| NVFP4 training, dense | VENDOR CLAIM: 2,520 PFLOPS, NVIDIA DGX page, as of 1 October 2026; 2.52 EFLOPS by our arithmetic, 2,520 / 1,000 |
| Aggregate NVLink bandwidth | 216 TB/s, NVIDIA datasheet as of 1 October 2026; 260 TB/s, NVIDIA SuperPOD blog dated 5 January 2026 |
| Cooling | 45°C warm-water liquid cooling; NVIDIA technical article, 16 March 2026 |
| Compute trays | Cable-free, hose-free and fanless design; NVIDIA technical article, 16 March 2026 |
| Network and switching chips | NVLink 6 Switch, ConnectX-9, BlueField-4 and Spectrum-6; NVIDIA platform announcement, 5 January 2026 |
The aggregate NVLink figures also disagree. Preserve their dates instead of building a new rack total from whichever per-GPU figure looks best. NVIDIA's January SuperPOD blog and the datasheet describe published configurations; neither resolves which specification a particular cloud allocation exposes.
NVIDIA's 5 January 2026 announcement defined six platform chip types, including Vera and Rubin. Its 16 March 2026 announcement expanded the platform to seven by adding the Groq 3 LPU. That wider platform scope matters when reading performance claims: a claim for a combined system should not become a claim for one Rubin GPU.
CoreWeave's 16 September 2026 announcement says its multi-rack cluster uses Spectrum-X Ethernet between Rubin racks. That is a separate connection from the rack's NVLink domain. Ask providers to describe both. The NVLink, PCIe and SXM guide explains the terminology; the GB200 and GB300 rack guide covers the preceding systems.
NVIDIA Rubin release date: manufacturing, shipment, access
The timeline preserves each publisher's status words. Read "in full production" as a manufacturing statement. Read "planned" as a target. For a rental decision, ask for an allocation and a start date after checking the announcement.
| Source date | Publisher | Exact status or dated plan |
|---|---|---|
| 2 June 2024 | NVIDIA | First revealed Rubin, paired with Vera on the roadmap. |
| 25 March 2025 | NVIDIA GTC recap | Previewed Vera Rubin under the former NVL144 name for the second half of 2026. |
| 5 January 2026 | NVIDIA | Said Rubin was in full production; partner products scheduled for the second half of 2026. |
| 5 January 2026 | Microsoft | Planned Rubin integration into Fairwater, including Wisconsin and Atlanta. |
| 5 January 2026 | Nebius | Targeted US and European service starting in H2 2026. |
| 16 March 2026 | NVIDIA | Expanded the platform; partner availability still scheduled to begin in the second half of 2026. |
| 16 March 2026 | Nebius | Meta agreement scheduled dedicated Vera Rubin capacity deliveries to begin in early 2027. |
| 16 March 2026 | Crusoe | Targeted deployments for late 2026 and throughout 2027. |
| 17 March 2026 | Oracle | Announced a Vera Rubin OCI Supercluster; no generally available instance launch date supplied. |
| April 2026 | Google Cloud | Announced A5X bare-metal instances based on Vera Rubin NVL72; scheduled for later in 2026. |
| 31 May 2026 | NVIDIA | Said the Vera Rubin platform was ramping into full production. |
| 1 June 2026 | CoreWeave | Successful bring-up and full-system validation of a rack. |
| 26 August 2026 | NVIDIA CFO Colette Kress | Production shipments commenced earlier that month. |
| 26 August 2026 | AWS | Planned additional Blackwell Ultra, Rubin and Rubin Ultra deployments across 2027 to 2028. |
| 27 August 2026 | NVIDIA | Reported AWS had received its first Vera CPU server and Vera Rubin GPU; the article announces no generally available Rubin EC2 instances. |
| 27 August 2026 | Tom's Hardware, attributing Satya Nadella | Reported Microsoft had deployed commercial Vera Rubin systems. |
| 16 September 2026 | CoreWeave | Announced a multi-rack Rubin cluster. |
| 30 September 2026 | CoreWeave / NVIDIA | CoreWeave announced cloud availability; NVIDIA described the offering as accessible to early-access customers. |
| As of 1 October 2026 | Lambda product page | Advertised H2 2026 availability through a sales-contact route. |
CoreWeave's 30 September 2026 release identifies Cognition as its first customer running production workloads on Vera Rubin. That is a concrete customer milestone. NVIDIA's same-day early-access description still matters: it does not establish unrestricted access for every buyer.
For Google, Oracle and Lambda, retain the announced or planned status shown above. For Microsoft and AWS, distinguish deployment or receipt of hardware from a public instance launch. Do not schedule a migration solely around a half-year target. Ask the provider to identify the service, region, access conditions and committed delivery date in writing.
NVIDIA Rubin price: reported rack costs, not a list price
ANALYST ESTIMATE: Tom's Hardware reported on 24 March 2026 that a VR200 NVL72 system would cost $5 million to $7 million, citing an anonymous source. The reported configuration included approximately $1 million of storage. Those numbers describe a system quotation with uncertain commercial terms, not a standalone accelerator price.
ANALYST ESTIMATE: Morgan Stanley, as reported by Tom's Hardware on 21 May 2026, estimated approximately $7.8 million per VR200 NVL72 rack paid by hyperscalers. The same report's body put a Rubin GPU at about $55,000 and a Vera CPU at about $5,000 inside volume rack purchases, although its headline says $50,000 per GPU; neither is a price for a part sold on its own. PC Gamer's 22 May 2026 account clarified that the rack estimate concerns customer acquisition cost, not NVIDIA's manufacturing cost. These reports do not establish a common configuration or a price trend.
Do not turn a rack acquisition estimate into your rental budget by dividing it by the GPU count. Obtain an actual service quote with the allocation, term, networking, storage and support specified. A hardware estimate does not contain those commercial terms.
VENDOR CLAIM: NVIDIA's 5 January 2026 launch release advertises up to 10 times lower inference cost per token than Blackwell. Its datasheet, as of 1 October 2026, attaches the one-tenth token-cost comparison to Kimi K2 Thinking with 32K input and 8K output, against GB200 NVL72. The chart omits precision. That omission limits how closely you can reproduce the comparison, and the result does not promise the same saving for your model.
Use that claim to frame an evaluation request. Ask for the model version, serving configuration, latency target and total billed resources. Compare the cost of completing the same work at the same quality target. Do not substitute a vendor's projected token economics for a provider's commercial terms.
What to rent while evaluating Rubin
Use the Rubin versus Blackwell decision guide to choose whether to migrate or keep the current deployment. Start with a repeatable workload and record its completion time, memory use and latency target. Require a Rubin evaluation to use the same model, precision and acceptance criteria before treating a headline throughput claim as a reason to move.
The live table below shows current per-GPU prices and stock for B200, B300, GB300 and H200; compare those rows against the complete allocation your job needs.
Use the GB300 rental page for listing details. Book the configuration that meets your deadline and workload requirements now. Move to Rubin when a provider commits to access and your workload evaluation justifies the full service cost.
Sources
- NVIDIA: first Rubin roadmap, 2 June 2024.
- NVIDIA: GTC roadmap recap, 25 March 2025.
- NVIDIA: rack branding editorial note, page dated 13 October 2025, as of 1 October 2026.
- SemiAnalysis: Rubin naming and operating configurations, 25 February 2026.
- NVIDIA: Rubin platform announcement, 5 January 2026.
- NVIDIA: Rubin technical overview, 5 January 2026.
- NVIDIA: Rubin GPU architecture, 21 July 2026.
- NVIDIA: Blackwell Ultra architecture, 22 August 2025.
- ServeTheHome: Rubin process and platform, 5 January 2026.
- NVIDIA: Vera Rubin NVL72 product specifications, as of 1 October 2026.
- NVIDIA: preliminary DGX Vera Rubin specifications, as of 1 October 2026.
- NVIDIA: Vera Rubin datasheet, as of 1 October 2026.
- NVIDIA: HGX systems and Blackwell reference, as of 1 October 2026 for Rubin; Blackwell Ultra entries in this site's GPU specification reference verified 13 September 2026.
- NVIDIA: preliminary DGX Rubin NVL8 specifications, as of 1 October 2026.
- NVIDIA: original CPX announcement, 9 September 2025.
- Tom's Hardware: Ian Buck press Q&A transcript, 23 March 2026.
- NVIDIA: Rubin rack topology and cooling, 16 March 2026.
- NVIDIA: DGX SuperPOD configurations, page dated 5 January 2026, as of 1 October 2026.
- NVIDIA: expanded Vera Rubin platform, 16 March 2026.
- NVIDIA: production ramp, 31 May 2026.
- NVIDIA: Colette Kress earnings-call remarks, 26 August 2026.
- CoreWeave: bring-up and full-system validation, 1 June 2026.
- CoreWeave: multi-rack cluster and Ethernet network, 16 September 2026.
- CoreWeave: cloud availability and Cognition deployment, 30 September 2026.
- NVIDIA: CoreWeave early access, 30 September 2026.
- Microsoft: Fairwater Rubin integration plans, 5 January 2026.
- Tom's Hardware: Microsoft deployment report, 27 August 2026.
- NVIDIA: AWS hardware delivery, updated 27 August 2026.
- AWS: multigeneration deployment commitment, 26 August 2026.
- Google Cloud: A5X announcement, April 2026.
- Oracle: Rubin Supercluster announcement, 17 March 2026.
- Nebius: US and Europe service plans, 5 January 2026.
- Nebius: dedicated Meta capacity plans, 16 March 2026.
- Lambda: Rubin service target, as of 1 October 2026.
- Crusoe: deployment plans, 16 March 2026.
- Tom's Hardware: reported rack acquisition estimates, 24 March 2026.
- Tom's Hardware: Morgan Stanley rack estimate, 21 May 2026.
- PC Gamer: acquisition-cost estimate clarification, 22 May 2026.