NVIDIA describes GB200 NVL72 as a liquid-cooled rack joining 36 Grace CPUs and 72 Blackwell GPUs into one NVLink domain. Its GB300 NVL72 uses Blackwell Ultra GPUs instead, with NVIDIA's rack figures raising GPU memory from 13.4 TB to 20 TB. Consider these racks when a single model, or the traffic between its GPUs, outgrows an 8-GPU server; otherwise, start your comparison with a smaller configuration.
The rest of this page works through what is actually in the rack, what NVIDIA means by "superchip," why the 72-GPU domain exists at all, and how to compare rental listings.
GB200 NVL72 vs GB300 NVL72 vs an 8-GPU B300 server
NVIDIA's rack specifications and DGX B300 documents supply the system figures below; the NVLink total for the 8-GPU platform comes from its HGX page. The purchase estimates have separate scopes and dates.
| Metric | GB200 NVL72 | GB300 NVL72 | 8-GPU DGX B300 |
|---|---|---|---|
| GPUs | 72 Blackwell | 72 Blackwell Ultra | 8 Blackwell Ultra |
| CPUs | 36 Grace | 36 Grace | 2 Intel Xeon 6776P |
| Total GPU memory | 13.4 TB HBM3e | 20 TB HBM3e | 2.3 TB at announcement; current page says 2.1 TB HBM3e |
| NVLink domain | 72 GPUs | 72 GPUs | 8 GPUs |
| Aggregate NVLink bandwidth | 130 TB/s | 130 TB/s | 14.4 TB/s |
| System power | Reported 120 to 140 kW; see caveat below | Reported 120 to 140 kW; see caveat below | 14.5 kW, 15 kW max |
| Cooling | liquid | liquid | air |
| FP4 (NVFP4) | 1,440 PFLOPS sparse, 720 dense | 1,440 PFLOPS sparse, 1,080 dense | 144 PFLOPS sparse, 108 dense (derived) |
| Purchase estimate (USD) | 3.1 million rack-only; 3.9 million with networking and storage (SemiAnalysis, 20 Aug 2025, ANALYST ESTIMATE) | 4.3 million (Wolfe Research, 30 Jan 2026, ANALYST ESTIMATE) | Lenovo example below is a different system |
NVIDIA publishes no list price for either rack, so these are dated analyst estimates. SemiAnalysis estimated on 20 August 2025 that a GB200 NVL72 rack costs about 3.1 million US dollars on its own or 3.9 million US dollars all in with networking and storage, an ANALYST ESTIMATE. Wolfe Research put GB200 NVL72 at about 3 million US dollars and GB300 NVL72 at about 4.3 million US dollars on 30 January 2026, also an ANALYST ESTIMATE.
For a separate 8-GPU purchase example, Lenovo's TCO paper gives a usual customer sale price of 785,606.50 US dollars for an 8x B300 ThinkSystem SR680a V4 as of 15 June 2026. That is a vendor-published sale price for a different system, not a DGX B300 quote. The table's 8-GPU FP4 figure is derived from the site's spec table: B300's per-GPU rating of 18 PFLOPS sparse and 13.5 PFLOPS dense, multiplied by eight GPUs. NVIDIA's 18 March 2025 announcement describes DGX B300 as air-cooled with 2.3 TB of GPU memory, while its product page recorded on 28 September 2026 says 2.1 TB. Check the exact configuration before comparing memory totals.
Tom's Hardware reported that power range on 12 August 2026. It is a press figure, not an NVIDIA rating, so size a facility from the vendor's documentation. NVIDIA describes both NVL72 racks as liquid-cooled; its air-cooled DGX B300 example does not establish that every 8-GPU server uses air cooling.
The superchip
NVIDIA's GB200 Superchip pairs one Grace CPU with two Blackwell GPUs, with 372 GB of HBM3e per superchip. Thirty-six of these superchips make up the 36 Grace CPUs and 72 Blackwell GPUs in one GB200 NVL72 rack.
NVIDIA's count for the GB300 NVL72 rack, 36 Grace CPUs and 72 Blackwell Ultra GPUs, gives the same one-CPU-to-two-GPU ratio (derived).
Do not confuse the rack products with DGX Station GB300: NVIDIA describes a desktop system pairing one Grace CPU with one Blackwell Ultra GPU over NVLink-C2C at 900 GB/s. It is a different product from the NVL72 rack.
Why a 72-GPU NVLink domain matters
NVIDIA's HGX page states that an 8-GPU HGX B200 board carries 14.4 TB/s of aggregate NVLink and 0.8 TB/s of networking, a derived 18.0 times ratio of the published totals. The two totals are not measured the same way, so read the ratio as scale, not speed. In multi-node clusters, traffic between separate NVLink domains uses a network fabric; this site covers InfiniBand and RoCE Ethernet in InfiniBand vs RoCE Ethernet. NVIDIA's GB200 and GB300 NVL72 specifications extend the domain to 72 GPUs, a derived 9.0 times the GPU count of an 8-GPU domain.
For a sense of scale, Tom's Hardware reported on 27 July 2024 that Meta's Llama 3 405 billion parameter training run used 16,384 H100 GPUs, with interruptions counted over a 54-day period. That example establishes the scale of a training cluster; it does not quantify how much of the same job's traffic would stay on NVLink in an NVL72 deployment.
Serving a large mixture-of-experts model can also involve communication across GPUs. DeepSeek's DeepEP library, used by vLLM and SGLang, provides all-to-all communication for expert parallelism. Its project benchmarks show a derived 8.1 to 12.1 times throughput gap between selected RDMA and NVLink rows. This is a VENDOR/PROJECT BENCHMARK comparing Hopper-class GPUs with ConnectX-7 NICs for RDMA against Blackwell-class GPUs for NVLink, not an isolated comparison of the two interconnects. SGLang's Mooncake transfer backend documents multi-node NVLink transport across NVL72 for key-value cache transfer. The mechanics of GPU communication are covered in NVLink vs PCIe vs SXM.
None of this makes a rack the default for smaller jobs. This site's NVIDIA DGX explained covers the 8-GPU DGX and HGX systems.
Specs, not rack totals
The table above describes whole systems. The table below is per GPU, pulled from the site's spec table, and it has no row for GB200. NVIDIA defines GB200 as a Grace CPU paired with two Blackwell GPUs; that does not make the B200 row a specification for the GB200 rack.
| Spec | GB300 | B300 | B200 |
|---|---|---|---|
| VRAM | 262 to 288 GB | 262 to 288 GB | 180 to 192 GB |
| Memory type | HBM3e | HBM3e | HBM3e |
| Memory bandwidth | 8,000 GB/s | 8,000 GB/s | 8,000 GB/s |
| FP8 (dense) | 4,500 TFLOPS | 4,500 TFLOPS | 4,500 TFLOPS |
| FP8 (with sparsity) | 9,000 TFLOPS | 9,000 TFLOPS | 9,000 TFLOPS |
| FP4 (dense) | 13,500 TFLOPS | 13,500 TFLOPS | 9,000 TFLOPS |
| FP4 (with sparsity) | 18,000 TFLOPS | 18,000 TFLOPS | 18,000 TFLOPS |
| TDP | 1,400 W | 1,400 W | 1,000 W |
| Interconnect | NVLink, PCIe 6.0, InfiniBand | NVLink, PCIe 6.0, InfiniBand | NVLink, PCIe 6.0, InfiniBand |
| Launch year | 2025 | 2025 | 2024 |
Read this table as each family's per-GPU numbers. Do not multiply them by 72 and assume they reproduce NVIDIA's rack specifications: the published rack FP4 totals above differ from that calculation. GB300 and B300 have matching family-level numeric fields in this site's table, but that alone does not establish identical configurations. Use the separately sourced system totals for the rack comparison. Full detail on the families in the table lives at the GPU spec reference.
Where to rent NVL72 capacity
NVIDIA does not sell racks directly to most buyers. Here is what each cloud below has said, in its own dated announcement, about when it turned on GB200 or GB300 NVL72 capacity:
- CoreWeave, a provider this site tracks live (see CoreWeave for its current offers), made GB200 NVL72 based instances generally available on 4 February 2025, a VENDOR CLAIM billing itself as the first cloud provider to do so. It named Cohere, IBM and Mistral AI as initial customers when it scaled the offering on 15 April 2025, then announced on 3 July 2025 what it called, again a VENDOR CLAIM, the first customer deployment of GB300 NVL72 systems.
- Google Cloud's own release notes date general availability of A4X, built on GB200 NVL72, to 10 September 2025. A4X Max, built on GB300 NVL72, began shipping in production on 28 October 2025, though Google's own post also calls that a preview, so treat production shipping and general availability as two different claims here.
- Oracle scheduled general availability of GB200 NVL72 based OCI Superclusters for April 2025 in a 26 March 2025 announcement, and its current Supercluster materials describe them as generally available. The same announcement made GB300 NVL72 Superclusters orderable, with general availability planned for later in 2025.
- Microsoft made Azure ND GB200 v6 virtual machines, each pairing two Grace CPUs with four Blackwell GPUs, generally available on 18 March 2025; eighteen of these VMs form one 72-GPU rack.
- AWS made Amazon EC2 P6e-GB200 UltraServers, which span 72 Blackwell GPUs in one NVLink domain, generally available on 9 July 2025, and P6e-GB300 UltraServers generally available on 2 December 2025; the P6e-GB300 launch notice directs customers to an AWS sales representative to get started.
- Nebius made GB200 NVL72 capacity generally available to European customers on 11 June 2025, and reported a live production GB300 NVL72 deployment at its expanded Finland data center on 17 December 2025.
- Crusoe added GB200 NVL72 instances at atNorth's ICE02 data center in Iceland on 25 August 2025, then on 15 September 2026 announced a multiyear Perplexity agreement that includes frontier model training on dedicated GB300 NVL72 clusters.
- Lambda announced on 3 November 2025 a multiyear deal to deploy AI infrastructure for Microsoft, including GB300 NVL72 systems.
- Together AI announced on 18 November 2024 a plan to co-build a 36,000 GPU GB200 NVL72 cluster with Hypertec Cloud starting in the first quarter of 2025. That remains an announced plan; the same page still directs interested customers to request early access rather than confirming a delivered cluster.
- Nscale reached NVIDIA Exemplar Cloud status on its GB300 NVL72 production fleet on 16 July 2026, a production milestone rather than an initial GA date.
None of these announcements state a rental price, and this site does not type one either. The tables below show what the GB300, B300 and B200 offers this site tracks cost right now.
| Variant | VRAM | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|---|
| GB300 SXM6 | 288 GB | none in stock | ||
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| GB300 SXM6 | none in stock | ||
| B300 SXM6 | $7.89 | RunPod | 1 |
| B200 SXM | $6.79 | RunPod | 2 |
Start at rent GB300 to check listings, use GB300 vs H100 for a comparison, or check GPU cluster capacity for larger configurations.
Who actually needs a rack
Consider a GB200 or GB300 NVL72 configuration if a single model does not fit inside one 8-GPU node's memory, or if expert-parallel traffic between GPUs is your bottleneck. Treat that as a reason to evaluate a larger NVLink domain, not proof that you need all 72 GPUs. vLLM's guidance recommends combining tensor and pipeline parallelism for models larger than a node; it does not prescribe a full rack.
Start with an 8-GPU B200 or B300 node if your model and its key-value cache fit inside one node. That follows vLLM's single-node guidance. Compare listings at rent B300 and rent B200.
Do not spend time on either rack if your workload fits on one GPU. Check the LLM VRAM calculator before assuming you need more than one card. Start with the smallest configuration that meets your memory and communication requirements.
NVIDIA's next rack-scale generation is Vera Rubin, built as Vera Rubin NVL72. NVIDIA's CFO Colette Kress said on 26 August 2026 that production shipments of Vera Rubin began earlier that month. Rubin vs Blackwell covers whether to wait for it, and the Vera Rubin explainer covers what changes in the rack.
Sources
- NVIDIA GB200 NVL72, recorded 28 September 2026
- NVIDIA GB300 NVL72, recorded 28 September 2026
- NVIDIA HGX platform, recorded 21 September 2026
- NVIDIA DGX Station, recorded 28 September 2026
- NVIDIA DGX B300 user guide, recorded 21 September 2026
- NVIDIA: Blackwell Ultra and DGX SuperPOD announcement, 18 March 2025
- Lenovo Press: on-premise vs cloud generative AI TCO, 2026 edition, updated 24 July 2026
- SemiAnalysis: H100 vs GB200 NVL72 training benchmarks, 20 August 2025
- Wolfe Research target note via Yahoo Finance, 30 January 2026
- Tom's Hardware: Vera Rubin NVL72 rack pricing, 24 March 2026
- Data Gravity: how much does an NVIDIA NVL72 cost, 31 August 2026
- Tom's Hardware: CoreWeave CEO on A100 contract into 2029, 12 August 2026
- Tom's Hardware: faulty H100 GPUs and HBM3 memory in Meta's Llama 3 training, 27 July 2024
- AWS EC2 P6 instance types, recorded 28 September 2026
- Microsoft Learn: ND GB200 v6 series, recorded 28 September 2026
- Lambda: multibillion-dollar agreement with Microsoft, 3 November 2025
- DeepSeek DeepEP repository, recorded 28 September 2026
- SGLang docs: prefill-decode disaggregation, recorded 28 September 2026
- CoreWeave: launches NVIDIA GB200 Grace Blackwell systems at scale, 15 April 2025
- Crusoe: Crusoe and Perplexity partnership, 15 September 2026
- vLLM: distributed inference guidance, recorded 21 September 2026
- NVIDIA DGX B300 product page, recorded 28 September 2026
- CoreWeave: first cloud provider with generally available GB200 NVL72 instances, 4 February 2025
- CoreWeave: first hyperscaler to deploy NVIDIA GB300 NVL72 platform, 3 July 2025
- Google Cloud AI Hypercomputer release notes, 10 September 2025
- Google Cloud: now shipping A4X Max, 28 October 2025
- Oracle: Supercluster with NVIDIA Blackwell, dedicated region and Alloy, 26 March 2025
- Oracle: accelerate AI with OCI Supercluster (PDF)
- Microsoft Azure: Microsoft and NVIDIA accelerate AI development and performance, 18 March 2025
- AWS: AI infrastructure with NVIDIA Blackwell, 9 July 2025
- AWS: Amazon EC2 P6e-GB300 UltraServers now generally available, 2 December 2025
- Nebius: first NVIDIA Blackwell general availability in Europe, 11 June 2025
- Nebius: first live NVIDIA GB300 NVL72 deployment in Europe, 17 December 2025
- Crusoe: expands Iceland data center capacity, 25 August 2025
- Together AI: NVIDIA GB200 and Together GPU cluster, 18 November 2024
- Nscale: achieves NVIDIA Exemplar Cloud status on GB300 NVL72, 16 July 2026
- NVIDIA Q2 fiscal 2027 earnings call transcript, 26 August 2026