NVIDIA Vera Rubin Explained: Specs, Release Date and Price

NVIDIA says Vera Rubin production shipments began in August 2026. Compare dated specs, conflicting bandwidth figures, cloud plans and reported rack costs.

By Faiz Ahmed•
•15 min read

NVIDIA Vera Rubin is the company's data center platform after Blackwell, pairing the Rubin GPU with the Vera CPU, as described in NVIDIA's 5 January 2026 announcement. Production shipments began in August 2026, CFO Colette Kress said on 26 August 2026. CoreWeave announced Vera Rubin NVL72 availability on its cloud on 30 September 2026; NVIDIA's post that day described the offering as accessible to early-access customers. NVL72 is the rack's current name; NVIDIA renamed it from NVL144, as dated in the next section.

As of 1 October 2026, this site's live price tables do not list Rubin yet, so the rent-now section shows Blackwell, with H200 alongside it. A shipment announcement is useful planning evidence. It does not tell you which allocation you can book, what software it includes or whether the provider will commit to your start date.

Decode the chip, superchip and rack names

NVIDIA's blog dated 13 October 2025 now carries this editorial note: "This blog has been updated to reflect a branding change from Vera Rubin NVL144 to Vera Rubin NVL72." That is the page's publication date, not a dated notice of when the edit happened. SemiAnalysis explained on 25 February 2026 that the old name counted 144 compute dies in 72 two-die GPU packages. The new name counts packages; SemiAnalysis dates the reversal to late December 2025.

All Vera Rubin NVL72 references below use that package-count convention. Do not treat the older name as evidence of twice as many GPUs. Keep the product name and its date together when reading an older rack proposal.

NameMeaning and dated attribution
VeraThe CPU paired with Rubin; NVIDIA, 5 January 2026. See the Vera CPU guide.
RubinThe GPU, with two compute dies unified in one package by NV-HBI; NVIDIA architecture article, 21 July 2026.
VR200Reported rack shorthand in Tom's Hardware's VR200 NVL72 coverage, 21 May 2026; not confirmation of a standalone GPU SKU.
Vera Rubin SuperchipTwo Rubin GPUs and one Vera CPU; NVIDIA datasheet, as of 1 October 2026.
Vera Rubin NVL72, formerly NVL144The rack with 72 Rubin GPUs and 36 Vera CPUs; NVIDIA DGX page, as of 1 October 2026. SemiAnalysis, 25 February 2026, says the 72 counts GPU packages; renaming explained above.
HGX Rubin NVL8Eight Rubin SXM GPUs on a baseboard; NVIDIA HGX page, as of 1 October 2026, distinguishes x86 and single-socket Vera versions.
DGX systems on RubinNVIDIA's complete DGX Vera Rubin NVL72 rack and DGX Rubin NVL8 system; NVIDIA product pages, as of 1 October 2026.
Rubin CPXContext-processing GPU announced by NVIDIA on 9 September 2025; Ian Buck later said NVIDIA had pulled it to prioritize LPU decode in 2026, in the transcript published 23 March 2026. See CPX and Groq LPX.
Rubin UltraA later generation, targeted for the second half of 2027 in NVIDIA's 25 March 2025 recap; a historical target, not a shipment statement. See the NVIDIA GPU roadmap.

For purchasing documents, write the full configuration. "NVIDIA VR200" alone leaves too much unspecified. The Superchip, an HGX baseboard and a complete DGX rack are different products. Ask for the GPU package count, CPU type, memory specification and network allocation in the same quote.

NVIDIA's HGX page, as of 1 October 2026, distinguishes HGX Vera Rubin NVL8 with a single-socket Vera CPU from an x86 variant whose host specifications belong to the OEM. NVIDIA's DGX Rubin NVL8 page, as of the same date, lists two Intel Xeon 6776P processors and marks its specifications preliminary. The Rubin name therefore does not identify the host CPU in every system. If your deployment requires a particular CPU environment, specify that requirement before comparing offers.

NVIDIA Rubin specs next to Blackwell Ultra

Both columns below are NVIDIA's own figures for one GPU package. The Rubin column comes from NVIDIA's 21 July 2026 architecture article and its Vera Rubin NVL72 product page, as of 1 October 2026. The Blackwell Ultra column comes from NVIDIA's 22 August 2025 Blackwell Ultra architecture article and the comparison table in its Rubin technical article, published 5 January 2026 and updated 16 March 2026. Compute figures are VENDOR CLAIM peak rates.

Per GPURubinBlackwell Ultra
Compute diesTwo, joined in one package by NV-HBITwo
Transistors336 billion208 billion
MemoryUp to 288 GB HBM4Up to 288 GB HBM3E
Memory bandwidthUp to 22 TB/s peak (July article); 19.2 TB/s (product page)8 TB/s
NVFP4, dense35 PFLOPS, labelled training15 PFLOPS
NVFP4, sparse50 PFLOPS, labelled inference20 PFLOPS
FP8, dense17.5 PFLOPS, FP8/FP6 training5 PFLOPS
NVLink per GPUNVLink 6: 3,600 GB/s (July article); 3 TB/s (product page)1,800 GB/s
NVLink-C2C to the CPU1.8 TB/s to Vera900 GB/s to Grace

By our arithmetic, Rubin has about 2.3 times the dense NVFP4, 2.5 times the sparse NVFP4 and 3.5 times the dense FP8 of Blackwell Ultra. VENDOR CLAIM: NVIDIA's July article calls the 22 TB/s figure 2.8 times Blackwell Ultra's 8 TB/s. These are ratios of peak specifications, not forecasts of how fast your job runs. Memory capacity does not grow: both top out at 288 GB per GPU, so the gain is bandwidth and compute, not room for a larger model.

Keep each comparison inside one source. This site's GPU specification reference lists the B300 at 13,500 dense and 18,000 sparse FP4 TFLOPS and 4,500 dense FP8 TFLOPS, figures taken from NVIDIA's HGX page. They differ from the architecture-article figures above, so never divide a Rubin number by one of them.

NVIDIA's two Rubin bandwidth pairs remain unresolved: 22 TB/s and 3,600 GB/s in the 21 July 2026 article, 19.2 TB/s and 3 TB/s on the product page as of 1 October 2026. Ask a provider which figure its delivered system meets rather than budgeting on the larger one.

Two more figures are reports, not NVIDIA specifications. ServeTheHome reported on 5 January 2026 that Rubin's compute dies are made on TSMC's 3nm process. ANALYST ESTIMATE: SemiAnalysis described Rubin operating configurations of 2,300 W (Max-P) and 1,800 W (Max-Q) on 25 February 2026. That is not a universal NVIDIA TDP, so take any hosting power figure from the system supplier.

For compute, keep all three labels: precision, sparsity and workload. NVIDIA's January technical article calls the 50 PFLOPS inference headline Transformer Engine compute and labels the 35 PFLOPS training figure dense; the product page labels the inference figure sparse. Do not halve one to get the other. The HBM4 guide covers the memory change in more detail.

NVIDIA's DGX page, as of 1 October 2026, lists 72 Rubin GPUs, 36 Vera CPUs and 20.7 TB HBM4 in DGX Vera Rubin NVL72. It explicitly marks the specifications preliminary and subject to change. NVIDIA's 16 March 2026 technical article describes the 72 GPUs as one scale-up domain connected through an NVLink copper spine.

NVIDIA's datasheet, as of 1 October 2026, also advertises up to 75 TB of fast-access memory, including CPU memory. That is not an HBM-only capacity. Keep GPU memory and host memory separate when sizing a model. The full rack total is also not a promise that your rental allocation includes the whole rack.

Rack itemPublished figure and qualification
NVFP4 inference, sparseVENDOR CLAIM: 3,600 PFLOPS, NVIDIA NVL72 page, as of 1 October 2026; 3.6 EFLOPS by our arithmetic, 3,600 / 1,000
NVFP4 training, denseVENDOR CLAIM: 2,520 PFLOPS, NVIDIA DGX page, as of 1 October 2026; 2.52 EFLOPS by our arithmetic, 2,520 / 1,000
Aggregate NVLink bandwidth216 TB/s, NVIDIA datasheet as of 1 October 2026; 260 TB/s, NVIDIA SuperPOD blog dated 5 January 2026
Cooling45°C warm-water liquid cooling; NVIDIA technical article, 16 March 2026
Compute traysCable-free, hose-free and fanless design; NVIDIA technical article, 16 March 2026
Network and switching chipsNVLink 6 Switch, ConnectX-9, BlueField-4 and Spectrum-6; NVIDIA platform announcement, 5 January 2026

The aggregate NVLink figures also disagree. Preserve their dates instead of building a new rack total from whichever per-GPU figure looks best. NVIDIA's January SuperPOD blog and the datasheet describe published configurations; neither resolves which specification a particular cloud allocation exposes.

NVIDIA's 5 January 2026 announcement defined six platform chip types, including Vera and Rubin. Its 16 March 2026 announcement expanded the platform to seven by adding the Groq 3 LPU. That wider platform scope matters when reading performance claims: a claim for a combined system should not become a claim for one Rubin GPU.

CoreWeave's 16 September 2026 announcement says its multi-rack cluster uses Spectrum-X Ethernet between Rubin racks. That is a separate connection from the rack's NVLink domain. Ask providers to describe both. The NVLink, PCIe and SXM guide explains the terminology; the GB200 and GB300 rack guide covers the preceding systems.

NVIDIA Rubin release date: manufacturing, shipment, access

The timeline preserves each publisher's status words. Read "in full production" as a manufacturing statement. Read "planned" as a target. For a rental decision, ask for an allocation and a start date after checking the announcement.

Source datePublisherExact status or dated plan
2 June 2024NVIDIAFirst revealed Rubin, paired with Vera on the roadmap.
25 March 2025NVIDIA GTC recapPreviewed Vera Rubin under the former NVL144 name for the second half of 2026.
5 January 2026NVIDIASaid Rubin was in full production; partner products scheduled for the second half of 2026.
5 January 2026MicrosoftPlanned Rubin integration into Fairwater, including Wisconsin and Atlanta.
5 January 2026NebiusTargeted US and European service starting in H2 2026.
16 March 2026NVIDIAExpanded the platform; partner availability still scheduled to begin in the second half of 2026.
16 March 2026NebiusMeta agreement scheduled dedicated Vera Rubin capacity deliveries to begin in early 2027.
16 March 2026CrusoeTargeted deployments for late 2026 and throughout 2027.
17 March 2026OracleAnnounced a Vera Rubin OCI Supercluster; no generally available instance launch date supplied.
April 2026Google CloudAnnounced A5X bare-metal instances based on Vera Rubin NVL72; scheduled for later in 2026.
31 May 2026NVIDIASaid the Vera Rubin platform was ramping into full production.
1 June 2026CoreWeaveSuccessful bring-up and full-system validation of a rack.
26 August 2026NVIDIA CFO Colette KressProduction shipments commenced earlier that month.
26 August 2026AWSPlanned additional Blackwell Ultra, Rubin and Rubin Ultra deployments across 2027 to 2028.
27 August 2026NVIDIAReported AWS had received its first Vera CPU server and Vera Rubin GPU; the article announces no generally available Rubin EC2 instances.
27 August 2026Tom's Hardware, attributing Satya NadellaReported Microsoft had deployed commercial Vera Rubin systems.
16 September 2026CoreWeaveAnnounced a multi-rack Rubin cluster.
30 September 2026CoreWeave / NVIDIACoreWeave announced cloud availability; NVIDIA described the offering as accessible to early-access customers.
As of 1 October 2026Lambda product pageAdvertised H2 2026 availability through a sales-contact route.

CoreWeave's 30 September 2026 release identifies Cognition as its first customer running production workloads on Vera Rubin. That is a concrete customer milestone. NVIDIA's same-day early-access description still matters: it does not establish unrestricted access for every buyer.

For Google, Oracle and Lambda, retain the announced or planned status shown above. For Microsoft and AWS, distinguish deployment or receipt of hardware from a public instance launch. Do not schedule a migration solely around a half-year target. Ask the provider to identify the service, region, access conditions and committed delivery date in writing.

NVIDIA Rubin price: reported rack costs, not a list price

ANALYST ESTIMATE: Tom's Hardware reported on 24 March 2026 that a VR200 NVL72 system would cost $5 million to $7 million, citing an anonymous source. The reported configuration included approximately $1 million of storage. Those numbers describe a system quotation with uncertain commercial terms, not a standalone accelerator price.

ANALYST ESTIMATE: Morgan Stanley, as reported by Tom's Hardware on 21 May 2026, estimated approximately $7.8 million per VR200 NVL72 rack paid by hyperscalers. The same report's body put a Rubin GPU at about $55,000 and a Vera CPU at about $5,000 inside volume rack purchases, although its headline says $50,000 per GPU; neither is a price for a part sold on its own. PC Gamer's 22 May 2026 account clarified that the rack estimate concerns customer acquisition cost, not NVIDIA's manufacturing cost. These reports do not establish a common configuration or a price trend.

Do not turn a rack acquisition estimate into your rental budget by dividing it by the GPU count. Obtain an actual service quote with the allocation, term, networking, storage and support specified. A hardware estimate does not contain those commercial terms.

VENDOR CLAIM: NVIDIA's 5 January 2026 launch release advertises up to 10 times lower inference cost per token than Blackwell. Its datasheet, as of 1 October 2026, attaches the one-tenth token-cost comparison to Kimi K2 Thinking with 32K input and 8K output, against GB200 NVL72. The chart omits precision. That omission limits how closely you can reproduce the comparison, and the result does not promise the same saving for your model.

Use that claim to frame an evaluation request. Ask for the model version, serving configuration, latency target and total billed resources. Compare the cost of completing the same work at the same quality target. Do not substitute a vendor's projected token economics for a provider's commercial terms.

What to rent while evaluating Rubin

Use the Rubin versus Blackwell decision guide to choose whether to migrate or keep the current deployment. Start with a repeatable workload and record its completion time, memory use and latency target. Require a Rubin evaluation to use the same model, precision and acceptance criteria before treating a headline throughput claim as a reason to move.

The live table below shows current per-GPU prices and stock for B200, B300, GB300 and H200; compare those rows against the complete allocation your job needs.

GPUCheapest $/GPU-hrProviderProviders in stock
B200$6.79RunPod3
B300$7.89RunPod3
GB300none in stock
H200$3.43QuantaCloud6
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Use the GB300 rental page for listing details. Book the configuration that meets your deadline and workload requirements now. Move to Rubin when a provider commits to access and your workload evaluation justifies the full service cost.

Sources

Frequently asked questions

What is NVIDIA Vera Rubin?▾

Vera Rubin is NVIDIA's data center platform after Blackwell, pairing the Rubin GPU with the Vera CPU. NVIDIA's January 2026 announcement also includes networking and switching chips in the platform.

What is the NVIDIA Rubin release date?▾

CFO Colette Kress said on 26 August 2026 that production shipments began earlier that month. CoreWeave announced cloud availability on 30 September 2026; NVIDIA's post that day described the offering as accessible to early-access customers.

Why did Vera Rubin NVL144 become NVL72?▾

NVIDIA's updated October 2025 blog acknowledges the branding change. SemiAnalysis explained on 25 February 2026 that NVL144 counted compute dies, while NVL72 counts the 72 GPU packages.

How much memory does a Rubin GPU have?▾

NVIDIA's 21 July 2026 architecture article specifies up to 288 GB HBM4. It gives up to 22 TB/s peak bandwidth, while NVIDIA's NVL72 product page, as of 1 October 2026, gives 19.2 TB/s.

Is VR200 an official standalone Rubin GPU name?▾

Tom's Hardware used VR200 NVL72 as rack shorthand on 21 May 2026. That usage does not establish a standalone GPU SKU; use Rubin GPU when specifying the accelerator.

Related Posts