GB200 and GB300 NVL72 Explained: Who Needs a Full Rack

What a GB200 or GB300 NVL72 rack is, how it compares with an 8-GPU B300 server, sourced purchase-price estimates, and how to compare rental listings.

By Faiz Ahmed•
•Updated 12 min read

NVIDIA describes GB200 NVL72 as a liquid-cooled rack joining 36 Grace CPUs and 72 Blackwell GPUs into one NVLink domain. Its GB300 NVL72 uses Blackwell Ultra GPUs instead, with NVIDIA's rack figures raising GPU memory from 13.4 TB to 20 TB. Consider these racks when a single model, or the traffic between its GPUs, outgrows an 8-GPU server; otherwise, start your comparison with a smaller configuration.

The rest of this page works through what is actually in the rack, what NVIDIA means by "superchip," why the 72-GPU domain exists at all, and how to compare rental listings.

GB200 NVL72 vs GB300 NVL72 vs an 8-GPU B300 server

NVIDIA's rack specifications and DGX B300 documents supply the system figures below; the NVLink total for the 8-GPU platform comes from its HGX page. The purchase estimates have separate scopes and dates.

MetricGB200 NVL72GB300 NVL728-GPU DGX B300
GPUs72 Blackwell72 Blackwell Ultra8 Blackwell Ultra
CPUs36 Grace36 Grace2 Intel Xeon 6776P
Total GPU memory13.4 TB HBM3e20 TB HBM3e2.3 TB at announcement; current page says 2.1 TB HBM3e
NVLink domain72 GPUs72 GPUs8 GPUs
Aggregate NVLink bandwidth130 TB/s130 TB/s14.4 TB/s
System powerReported 120 to 140 kW; see caveat belowReported 120 to 140 kW; see caveat below14.5 kW, 15 kW max
Coolingliquidliquidair
FP4 (NVFP4)1,440 PFLOPS sparse, 720 dense1,440 PFLOPS sparse, 1,080 dense144 PFLOPS sparse, 108 dense (derived)
Purchase estimate (USD)3.1 million rack-only; 3.9 million with networking and storage (SemiAnalysis, 20 Aug 2025, ANALYST ESTIMATE)4.3 million (Wolfe Research, 30 Jan 2026, ANALYST ESTIMATE)Lenovo example below is a different system

NVIDIA publishes no list price for either rack, so these are dated analyst estimates. SemiAnalysis estimated on 20 August 2025 that a GB200 NVL72 rack costs about 3.1 million US dollars on its own or 3.9 million US dollars all in with networking and storage, an ANALYST ESTIMATE. Wolfe Research put GB200 NVL72 at about 3 million US dollars and GB300 NVL72 at about 4.3 million US dollars on 30 January 2026, also an ANALYST ESTIMATE.

For a separate 8-GPU purchase example, Lenovo's TCO paper gives a usual customer sale price of 785,606.50 US dollars for an 8x B300 ThinkSystem SR680a V4 as of 15 June 2026. That is a vendor-published sale price for a different system, not a DGX B300 quote. The table's 8-GPU FP4 figure is derived from the site's spec table: B300's per-GPU rating of 18 PFLOPS sparse and 13.5 PFLOPS dense, multiplied by eight GPUs. NVIDIA's 18 March 2025 announcement describes DGX B300 as air-cooled with 2.3 TB of GPU memory, while its product page recorded on 28 September 2026 says 2.1 TB. Check the exact configuration before comparing memory totals.

Tom's Hardware reported that power range on 12 August 2026. It is a press figure, not an NVIDIA rating, so size a facility from the vendor's documentation. NVIDIA describes both NVL72 racks as liquid-cooled; its air-cooled DGX B300 example does not establish that every 8-GPU server uses air cooling.

The superchip

NVIDIA's GB200 Superchip pairs one Grace CPU with two Blackwell GPUs, with 372 GB of HBM3e per superchip. Thirty-six of these superchips make up the 36 Grace CPUs and 72 Blackwell GPUs in one GB200 NVL72 rack.

NVIDIA's count for the GB300 NVL72 rack, 36 Grace CPUs and 72 Blackwell Ultra GPUs, gives the same one-CPU-to-two-GPU ratio (derived).

Do not confuse the rack products with DGX Station GB300: NVIDIA describes a desktop system pairing one Grace CPU with one Blackwell Ultra GPU over NVLink-C2C at 900 GB/s. It is a different product from the NVL72 rack.

NVIDIA's HGX page states that an 8-GPU HGX B200 board carries 14.4 TB/s of aggregate NVLink and 0.8 TB/s of networking, a derived 18.0 times ratio of the published totals. The two totals are not measured the same way, so read the ratio as scale, not speed. In multi-node clusters, traffic between separate NVLink domains uses a network fabric; this site covers InfiniBand and RoCE Ethernet in InfiniBand vs RoCE Ethernet. NVIDIA's GB200 and GB300 NVL72 specifications extend the domain to 72 GPUs, a derived 9.0 times the GPU count of an 8-GPU domain.

For a sense of scale, Tom's Hardware reported on 27 July 2024 that Meta's Llama 3 405 billion parameter training run used 16,384 H100 GPUs, with interruptions counted over a 54-day period. That example establishes the scale of a training cluster; it does not quantify how much of the same job's traffic would stay on NVLink in an NVL72 deployment.

Serving a large mixture-of-experts model can also involve communication across GPUs. DeepSeek's DeepEP library, used by vLLM and SGLang, provides all-to-all communication for expert parallelism. Its project benchmarks show a derived 8.1 to 12.1 times throughput gap between selected RDMA and NVLink rows. This is a VENDOR/PROJECT BENCHMARK comparing Hopper-class GPUs with ConnectX-7 NICs for RDMA against Blackwell-class GPUs for NVLink, not an isolated comparison of the two interconnects. SGLang's Mooncake transfer backend documents multi-node NVLink transport across NVL72 for key-value cache transfer. The mechanics of GPU communication are covered in NVLink vs PCIe vs SXM.

None of this makes a rack the default for smaller jobs. This site's NVIDIA DGX explained covers the 8-GPU DGX and HGX systems.

Specs, not rack totals

The table above describes whole systems. The table below is per GPU, pulled from the site's spec table, and it has no row for GB200. NVIDIA defines GB200 as a Grace CPU paired with two Blackwell GPUs; that does not make the B200 row a specification for the GB200 rack.

SpecGB300B300B200
VRAM262 to 288 GB262 to 288 GB180 to 192 GB
Memory typeHBM3eHBM3eHBM3e
Memory bandwidth8,000 GB/s8,000 GB/s8,000 GB/s
FP8 (dense)4,500 TFLOPS4,500 TFLOPS4,500 TFLOPS
FP8 (with sparsity)9,000 TFLOPS9,000 TFLOPS9,000 TFLOPS
FP4 (dense)13,500 TFLOPS13,500 TFLOPS9,000 TFLOPS
FP4 (with sparsity)18,000 TFLOPS18,000 TFLOPS18,000 TFLOPS
TDP1,400 W1,400 W1,000 W
InterconnectNVLink, PCIe 6.0, InfiniBandNVLink, PCIe 6.0, InfiniBandNVLink, PCIe 6.0, InfiniBand
Launch year202520252024
Figures from the vendor datasheets: GB300, B300, B200, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

Read this table as each family's per-GPU numbers. Do not multiply them by 72 and assume they reproduce NVIDIA's rack specifications: the published rack FP4 totals above differ from that calculation. GB300 and B300 have matching family-level numeric fields in this site's table, but that alone does not establish identical configurations. Use the separately sourced system totals for the rack comparison. Full detail on the families in the table lives at the GPU spec reference.

Where to rent NVL72 capacity

NVIDIA does not sell racks directly to most buyers. Here is what each cloud below has said, in its own dated announcement, about when it turned on GB200 or GB300 NVL72 capacity:

  • CoreWeave, a provider this site tracks live (see CoreWeave for its current offers), made GB200 NVL72 based instances generally available on 4 February 2025, a VENDOR CLAIM billing itself as the first cloud provider to do so. It named Cohere, IBM and Mistral AI as initial customers when it scaled the offering on 15 April 2025, then announced on 3 July 2025 what it called, again a VENDOR CLAIM, the first customer deployment of GB300 NVL72 systems.
  • Google Cloud's own release notes date general availability of A4X, built on GB200 NVL72, to 10 September 2025. A4X Max, built on GB300 NVL72, began shipping in production on 28 October 2025, though Google's own post also calls that a preview, so treat production shipping and general availability as two different claims here.
  • Oracle scheduled general availability of GB200 NVL72 based OCI Superclusters for April 2025 in a 26 March 2025 announcement, and its current Supercluster materials describe them as generally available. The same announcement made GB300 NVL72 Superclusters orderable, with general availability planned for later in 2025.
  • Microsoft made Azure ND GB200 v6 virtual machines, each pairing two Grace CPUs with four Blackwell GPUs, generally available on 18 March 2025; eighteen of these VMs form one 72-GPU rack.
  • AWS made Amazon EC2 P6e-GB200 UltraServers, which span 72 Blackwell GPUs in one NVLink domain, generally available on 9 July 2025, and P6e-GB300 UltraServers generally available on 2 December 2025; the P6e-GB300 launch notice directs customers to an AWS sales representative to get started.
  • Nebius made GB200 NVL72 capacity generally available to European customers on 11 June 2025, and reported a live production GB300 NVL72 deployment at its expanded Finland data center on 17 December 2025.
  • Crusoe added GB200 NVL72 instances at atNorth's ICE02 data center in Iceland on 25 August 2025, then on 15 September 2026 announced a multiyear Perplexity agreement that includes frontier model training on dedicated GB300 NVL72 clusters.
  • Lambda announced on 3 November 2025 a multiyear deal to deploy AI infrastructure for Microsoft, including GB300 NVL72 systems.
  • Together AI announced on 18 November 2024 a plan to co-build a 36,000 GPU GB200 NVL72 cluster with Hypertec Cloud starting in the first quarter of 2025. That remains an announced plan; the same page still directs interested customers to request early access rather than confirming a delivered cluster.
  • Nscale reached NVIDIA Exemplar Cloud status on its GB300 NVL72 production fleet on 16 July 2026, a production milestone rather than an initial GA date.

None of these announcements state a rental price, and this site does not type one either. The tables below show what the GB300, B300 and B200 offers this site tracks cost right now.

VariantVRAMCheapest $/GPU-hrProviderProviders in stock
GB300 SXM6288 GBnone in stock
Cheapest in-stock on-demand price per GPU-hour for each GB300 variant, from providers with live stock tracking.
GPUCheapest $/GPU-hrProviderProviders in stock
GB300 SXM6none in stock
B300 SXM6$7.89RunPod1
B200 SXM$6.79RunPod2
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

Start at rent GB300 to check listings, use GB300 vs H100 for a comparison, or check GPU cluster capacity for larger configurations.

Who actually needs a rack

Consider a GB200 or GB300 NVL72 configuration if a single model does not fit inside one 8-GPU node's memory, or if expert-parallel traffic between GPUs is your bottleneck. Treat that as a reason to evaluate a larger NVLink domain, not proof that you need all 72 GPUs. vLLM's guidance recommends combining tensor and pipeline parallelism for models larger than a node; it does not prescribe a full rack.

Start with an 8-GPU B200 or B300 node if your model and its key-value cache fit inside one node. That follows vLLM's single-node guidance. Compare listings at rent B300 and rent B200.

Do not spend time on either rack if your workload fits on one GPU. Check the LLM VRAM calculator before assuming you need more than one card. Start with the smallest configuration that meets your memory and communication requirements.

NVIDIA's next rack-scale generation is Vera Rubin, built as Vera Rubin NVL72. NVIDIA's CFO Colette Kress said on 26 August 2026 that production shipments of Vera Rubin began earlier that month. Rubin vs Blackwell covers whether to wait for it, and the Vera Rubin explainer covers what changes in the rack.

Sources

Frequently asked questions

What is the difference between GB200 NVL72 and GB300 NVL72?▾

NVIDIA describes both as connecting 36 Grace CPUs and 72 Blackwell family GPUs in one NVLink domain. Its GB300 Blackwell Ultra figures give 20 TB of rack GPU memory against GB200's 13.4 TB, and 1,080 PFLOPS of dense FP4 against 720 PFLOPS, while peak sparse FP4 and NVLink bandwidth stay the same.

What does GB200 mean?▾

GB200 is NVIDIA's name for a Grace Blackwell Superchip: one Grace CPU paired with two Blackwell GPUs. NVIDIA's rack configuration contains thirty-six of these superchips, totaling 36 Grace CPUs and 72 GPUs.

Can I rent a slice of a GB200 or GB300 NVL72 rack instead of buying one?▾

Yes, in smaller units. Microsoft's ND GB200 v6 virtual machine has four Blackwell GPUs, and eighteen of them form one 72-GPU rack. The live tables on this page show the GB300 offers the site tracks.

How much does a GB200 or GB300 NVL72 rack cost?▾

SemiAnalysis's 20 August 2025 ANALYST ESTIMATE put a GB200 NVL72 rack at 3.1 million US dollars alone or 3.9 million with networking and storage; Wolfe Research's 30 January 2026 ANALYST ESTIMATE put GB300 NVL72 at about 4.3 million US dollars. These are estimates, not NVIDIA list prices.

Do I need a 72-GPU NVLink domain for LLM inference?▾

Consider it when model memory or communication between GPUs makes one 8-GPU node unsuitable. SGLang documents multi-node NVLink transport across NVL72 for key-value cache transfer, but that does not make a full rack necessary for every such workload.

Why is GB200 missing from the site's GPU spec table?▾

NVIDIA defines a GB200 Superchip as one Grace CPU paired with two Blackwell GPUs. The site has no separate GB200 spec-table entry; use the GB300, B300 and B200 entries for their respective per-GPU figures, and the system table for separately sourced rack totals.

Related Posts