Rubin vs Blackwell: Rent Now, Keep an Exit Clause

Compare live Blackwell costs with dated Rubin rollout plans, vendor performance claims and contract terms so you can choose whether to rent, reserve or wait.

By Faiz Ahmed•
•12 min read

For work starting in the next few months, rent Blackwell: NVIDIA CFO Colette Kress said on 26 August 2026 that Vera Rubin production shipments began that month, and CoreWeave's 30 September availability announcement was described by NVIDIA that day as early access. Other cloud plans target H2 2026 or 2027, with Nebius's 16 March 2026 Meta agreement scheduling dedicated Vera Rubin capacity deliveries from early 2027. Change that answer only when you have a written Rubin allocation that meets your deadline and a workload trial that justifies waiting.

Your decision is about a start date, useful throughput and the cost of committing. A future system, however fast, does not finish a job during the months you spend waiting for it. Equally, a long reservation deserves more scrutiny than a short rental. Keep the right to change hardware when the benefit becomes measurable.

Price the work you can start now

Use the live table to establish your Blackwell budget. It also includes Hopper options as a check on whether your workload needs the newer family at all. As of 1 October 2026, this site's live tables do not list Rubin.

GPUCheapest $/GPU-hrProviderProviders in stock
H100$2.59QuantaCloud8
H200$3.43QuantaCloud7
B200$7.20VERDA1
B300$7.89RunPod1
GB300none in stock
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.
Pricing typeCheapest listed $/GPU-hrvs on-demandProviderListingListings seen
On-demand$4.36baselineCoreWeave8 GPUs7 from 3 providers
Spot or preemptible$3.6017% lowerVERDA2 GPUs1 from 1 provider
Reserved$8.87103% higherLambda Labs1536 GPUs, 1-week minimum term16 from 1 provider
Listed prices for B200: the cheapest catalogue listing of each pricing type seen in the last 24 hours, per GPU-hour, from secure (non-P2P) providers. Listed prices are not checked for stock, so they can differ from the in-stock on-demand price. Latest listing observation: .

Compare reserved with on-demand B200 pricing before committing now, and require an exit or migration clause if the reservation extends into your planned Rubin evaluation.

GPU$/GPU-hrPer day (×24)Per month (×730)Provider
B200$7.20$172.80$5,256VERDA
B300$7.89$189.36$5,760RunPod
H200$3.43$82.32$2,504QuantaCloud
Cheapest in-stock on-demand price per GPU-hour, for one GPU running the whole time. Per day = hourly price × 24 hours. Per month = hourly price × 730 hours (365 days × 24 ÷ 12), rounded to the nearest dollar. Storage, data transfer and tax are not included. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Start with the B200 listings, B300 listings or GB300 listings. Use the GPU price index to revisit the comparison before signing. The monthly table is a budgeting input, not a substitute for a quote covering your exact allocation.

Write down the amount of work you need completed, the deadline and your acceptable failure rate. Then ask each bidder for the same workload trial. Include storage, network transfers, idle allocation and migration time in your budget. A low unit cost is not enough if the contract forces you to pay for unused capacity. If buying is also on the table, use the rent-versus-buy calculator with your own hardware quote and utilization assumptions.

Rubin vs Blackwell specifications without mixed baselines

The B200 and B300 figures below come from this site's specification table, verified against NVIDIA's HGX page on 13 September 2026. Rubin figures come from NVIDIA's Vera Rubin NVL72 product page, as of 1 October 2026. Compute figures are peak VENDOR CLAIM values; the precision and sparsity labels matter more than the largest number.

Per GPUB200B300Rubin, NVIDIA page as of 1 October 2026
Memory capacity180 to 192 GB262 to 288 GB288 GB
Memory typeHBM3eHBM3eHBM4
Memory bandwidth8,000 GB/s8,000 GB/s19.2 TB/s
FP4 dense9,000 TFLOPS13,500 TFLOPSNVFP4 training: 35 PFLOPS, dense
FP4 sparse18,000 TFLOPS18,000 TFLOPSNVFP4 inference: 50 PFLOPS, sparse
FP8 dense4,500 TFLOPS4,500 TFLOPSFP8/FP6 training: 17.5 PFLOPS, dense

These are family specifications, not a promise about every cloud instance. Confirm the memory exposed by the quoted configuration. The B300 versus B200 comparison helps narrow that choice before you negotiate a longer commitment.

NVIDIA's 21 July 2026 architecture article gives Rubin up to 22 TB/s peak HBM4 bandwidth and 3,600 GB/s NVLink bandwidth. NVIDIA's product page, as of 1 October 2026, instead gives 19.2 TB/s and 3 TB/s. That disagreement is unresolved here. Ask the provider which specification its delivered system meets rather than building a budget around the larger figures.

Do not divide across those columns: the Rubin rows separate inference from training, while the B200 and B300 columns keep the specification table's labels. NVIDIA's own pages do support one like-for-like comparison. Its 22 August 2025 Blackwell Ultra architecture article gives 15 PFLOPS dense NVFP4 and 5 PFLOPS dense FP8 per GPU; against Rubin's 35 and 17.5 PFLOPS dense, that is about 2.3 and 3.5 times by our arithmetic. Both sides are VENDOR CLAIM peak rates, not job results. The B300 column's lower figures, 13,500 dense FP4 and 4,500 dense FP8 TFLOPS, come from NVIDIA's HGX page, a different source, so keep the two sets apart. Also notice that Rubin's listed memory capacity matches the B300 family's maximum. Waiting solely to fit a larger model on one GPU needs a more specific justification.

For a search such as "VR200 vs GB300," keep the product names straight. Tom's Hardware used VR200 NVL72 as rack shorthand on 21 May 2026. NVIDIA's official names in its 5 January announcement are Rubin GPU, Vera CPU and Vera Rubin platform. SemiAnalysis explained on 25 February 2026 that the renamed Vera Rubin NVL72 counts 72 GPU packages; the earlier NVL144 name counted 144 compute dies. The Vera Rubin platform guide covers the system, while the Blackwell rack guide explains the GB200 and GB300 rental context.

What the token and training claims mean for your job

NVIDIA's 5 January 2026 announcement advertises up to 10× lower inference cost per token than Blackwell, a VENDOR CLAIM. NVIDIA's datasheet, as of 1 October 2026, makes the same VENDOR CLAIM more specific: one-tenth token cost for Kimi K2 Thinking with 32K input and 8K output, comparing Rubin NVL72 with GB200 NVL72. That chart omits precision. It is not a quoted reduction in your rental bill.

For Vera Rubin vs GB300, that baseline is especially important. The datasheet chart does not establish the same gain over GB300. Kress's 26 August 2026 statement instead claims 35× lower token cost for the expanded Vera Rubin platform than Grace Blackwell Ultra, another VENDOR CLAIM. The call gives no precision or workload protocol. Do not substitute that broader platform claim for a measured result on the GPU configuration you plan to rent.

For training, NVIDIA's product page, as of 1 October 2026, claims 35 PFLOPS dense NVFP4 and 17.5 PFLOPS dense FP8/FP6 per Rubin GPU. Those are VENDOR CLAIM peak rates, not workload results. They do not tell you how many GPUs your training run can eliminate or when it will finish.

The strongest reason to evaluate Rubin is a workload whose cost is dominated by the operations NVIDIA targets: sustained inference or substantial training compute. Treat that as a trial hypothesis. Measure completed requests at your latency and quality target, or training progress to your chosen checkpoint. Keep the model, precision, context length and acceptance criteria fixed between bids. Otherwise, an apparent hardware gain may come from changing the job.

Keep the software plan concrete

Do not budget on an unqualified promise that every CUDA application will move unchanged. Put your container, framework build and custom kernels into the acceptance test, and ask the provider to identify its supported stack before you commit.

NVFP4 is a point of continuity: NVIDIA's Blackwell Ultra architecture article of 22 August 2025 and its Rubin product page, as of 1 October 2026, both specify it. NVIDIA's Rubin technical article, published 5 January 2026 and updated 16 March 2026, says the third-generation Transformer Engine adds hardware-accelerated adaptive compression for NVFP4. That is a hardware feature, not evidence that your current execution path automatically uses it.

NVIDIA's DGX Vera Rubin page, as of 1 October 2026, lists Mission Control, AI Enterprise and DGX OS in the software bundle. Ask what the cloud contract actually includes. Keep your migration estimate separate from the hardware comparison until your application passes its own correctness and performance checks.

The facility determines who can host the rack

NVIDIA's 16 March 2026 Vera Rubin POD article describes cable-free, hose-free and fanless compute trays, 45°C warm-water liquid cooling, and rack-level energy storage for Intelligent Power Smoothing. This is a facility and rack integration decision as well as a chip upgrade.

Supermicro's Vera Rubin page, as of 1 October 2026, sizes its NVL72 DLC-2 cooling design for 227 kW per rack. That is an OEM cooling-design figure, not measured workload consumption or a universal Rubin rack power rating. NVIDIA's DGX Vera Rubin specifications, as of the same date, remain preliminary and subject to change.

A sound procurement rule: require evidence that the host has commissioned the cooling, power and rack configuration being quoted. A deployment target should not carry the same weight as an accepted allocation. Ask who bears the cost if facility readiness delays your start. For a renter, that contract term is more actionable than estimating a rack's electricity bill from a cooling figure.

When renters might see Rubin

The status table preserves the dated announcement language. Read each row as product history or a stated plan, not as a statement of spare capacity today.

CloudDated publisher statementStatus and target
CoreWeaveCoreWeave, 30 September 2026; NVIDIA, same dayCoreWeave announced availability of Vera Rubin NVL72, naming Cognition as its first production customer; NVIDIA described access for early-access customers.
Google Cloud A5XGoogle, April 2026Announced A5X bare-metal instances based on Vera Rubin NVL72, scheduled for later in 2026; not a general-availability announcement.
NebiusNebius, 5 January 2026Targeted US and Europe service starting in H2 2026, through AI Cloud and Token Factory.
LambdaLambda's page, as of 1 October 2026Advertises H2 2026 availability through a sales-contact route, rather than a dated general-availability announcement.
CrusoeCrusoe, 16 March 2026Targets Rubin GPU/NVL72 deployments for late 2026 and throughout 2027.
OracleOracle, 17 March 2026Announced a next-generation OCI Supercluster based on Vera Rubin; no generally available instance launch date supplied.
MicrosoftTom's Hardware, 27 August 2026Reported deployment of commercial Vera Rubin systems, attributing the milestone to Satya Nadella. Deployment is not a public booking commitment.
AWSAWS, 26 August 2026; NVIDIA, 27 August 2026AWS plans additional Blackwell Ultra, Rubin and Rubin Ultra GPUs across 2027 to 2028; NVIDIA said AWS received its first Vera CPU server and Vera Rubin GPU. No generally available Rubin EC2 launch announced by that delivery article.

Some announced capacity has a named destination. Nebius's 16 March 2026 Meta agreement schedules dedicated Vera Rubin capacity deliveries to begin in early 2027. NVIDIA's 10 March 2026 Thinking Machines Lab partnership targets at least 1 GW of Vera Rubin deployments beginning in early 2027. Neither announcement books capacity for your team.

If you are tracking NVIDIA after Blackwell, use the GPU roadmap for the longer sequence. For this purchase, insist on a region, allocation size, start date and acceptance process. Do not let a roadmap date stand in for those terms.

Choose the commitment that matches your deadline

Workload or situationDecisionReason
Short project starting nowRent Blackwell nowFinish useful work before spending time on an unallocated future system.
Multi-month training runReserve with an exit clauseSecure continuity while protecting against delayed starts or a justified hardware change.
Inference serving with a one-year reservationReserve with an exit clauseCompare delivered serving cost, then negotiate migration rights before locking the term.
Waiting for a quoted Rubin contractWaitOnly if the written allocation meets your deadline and its acceptance trial supports the economics.
Small team needing one or eight GPUsRent Blackwell nowEvaluate the allocation you need; a rack announcement does not settle your booking.

If your current baseline is Hopper, the GB300 versus H100 comparison can frame the intermediate upgrade. You do not have to make the entire hardware roadmap part of one contract.

Rent Blackwell for the next job. For a longer commitment, negotiate an exit before signing. Wait for Rubin only when the allocation and workload result are concrete enough to replace a working plan.

Sources

Frequently asked questions

Should I wait for Rubin before renting Blackwell?▾

Rent Blackwell for work starting in the next few months. Wait only when a written Rubin allocation meets your deadline and a workload trial supports the total cost.

Does Rubin's token-cost claim apply to GB300?▾

NVIDIA's one-tenth token-cost chart, a VENDOR CLAIM, compares Rubin with GB200, not GB300. Its separate August claim, also a VENDOR CLAIM, uses Grace Blackwell Ultra but lacks a workload and precision protocol.

Does Rubin have more memory than B300?▾

NVIDIA lists 288 GB HBM4 for Rubin, matching the B300 family's maximum capacity in our specification table. Bandwidth and memory generation change, so capacity alone does not settle the choice.

Can I treat a Rubin cloud announcement as a booking?▾

No. CoreWeave announced availability on 30 September 2026, while NVIDIA described early access that day; other announcements include future targets. Get an allocation, start date and acceptance terms in writing.

Should I sign a one-year Blackwell reservation?▾

Reserve with an exit clause if the workload justifies the commitment. Compare any discount against migration costs and the value of being able to leave.

Related Posts