ReferenceUpdated September 2026Live reserved cluster rates16–1,024+ GPUs

GPU Cluster Pricing, 2026

A GPU cluster is more than one GPU node — typically 8 GPUs each, linked inside the node by NVLink — joined by a node-to-node fabric such as InfiniBand, RoCEv2 or high-speed Ethernet so they train and serve as one machine. This page covers what a cluster costs in 2026, how reserved cluster capacity compares to renting single nodes on-demand, and when each makes sense.

As a single-node baseline: an 8× H100 node on-demand runs roughly $15–28/hr, an 8× H200 node $24–36/hr, and an 8× A100 node $9–18/hr. Those are one node with no fabric between nodes. Multi-node reserved clusters on 3–12 month terms typically price 25–45% below on-demand per GPU-hour.

Live today: Lambda Labs reserved H100 SXM5 clusters from $5.54/GPU-hr and Lambda Labs reserved B200 SXM clusters from $8.87/GPU-hr, 16 to 1,536 GPUs on InfiniBand with a 1-week minimum. See the live cluster rates.

What Is a GPU Cluster?

A single node is one server, almost always 8 GPUs connected internally by NVLink. A cluster is two or more nodes connected by a node-to-node fabric — InfiniBand, RoCEv2 over Ethernet, or plain high-speed Ethernet — so dozens or thousands of GPUs cooperate on a single training or inference workload. NVLink never leaves the node; the fabric is what makes it a cluster.

Clusters are how teams pre-train and large-fine-tune models that don't fit on one node, and how high-throughput inference fleets are run with predictable capacity. The defining feature versus on-demand single GPUs is the interconnect: it keeps every GPU fed during the all-reduce communication that dominates distributed training.

Live GPU Cluster Pricing

Reserved multi-node cluster rates published by providers, verified against their pricing pages and refreshed hourly. Per-GPU-hour pricing steps down with cluster size; every listing names its node-to-node fabric and minimum term. Pick a size in the live table to compare it against on-demand nodes.

Lambda Labs · 16512× H100 SXM5

8-GPU nodes · 80 GB VRAM per GPU · NVLink within each node · InfiniBand between nodes · 1-week minimum

$5.54/GPU/hr
from $44.32/node-hr
Cluster sizePer GPU-hour
16–63 GPUs$6.16
64–255 GPUs$5.85
256+ GPUs$5.54
Sizes: 16, 32, 64, 128, 256, 512 GPUs·Regions: Texas, US, Washington DC, US·Per node: 208 vCPU, 1,800 GB RAM, 22 TB SSD

Lambda Labs · 161536× B200 SXM

8-GPU nodes · 192 GB VRAM per GPU · NVLink within each node · InfiniBand between nodes · 1-week minimum

$8.87/GPU/hr
from $70.96/node-hr
Cluster sizePer GPU-hour
16–63 GPUs$9.86
64–255 GPUs$9.36
256+ GPUs$8.87
Sizes: 16, 32, 64, 128, 256, 512, 1024, 1536 GPUs·Regions: Ohio, US, Texas, US·Per node: 208 vCPU, 2,900 GB RAM, 22 TB SSD

On-Demand Node Baseline

Not cluster prices: what a single 8-GPU node costs by the hour, with no fabric between nodes. This is the number a cluster quote is compared against. Per-GPU ranges are typical 2026 market rates; the cheapest live 8-GPU node comes from our tracker and refreshes hourly.

GPUOn-Demand / GPU·hrCheapest 8-GPU Node Now
H100 SXM 80GB$1.90–3.50$19.51/hrCoreWeave · United States · $2.44/GPU
H200 SXM 141GB$3.00–4.50$20.64/hrCoreWeave · United States · $2.58/GPU
B200 SXM$4.00–6.50$34.87/hrCoreWeave · United States · $4.36/GPU
GB300 NVL$5.50–9.00Rack-scale (NVL72)typical range
A100 SXM 80GB$1.10–2.20$9.20/hrDenvr · Virginia · $1.15/GPU
MI300X$1.90–3.50$24.64/hrCirrascale · United States · $3.08/GPU

Single-node on-demand, not cluster quotes. Per-GPU rates update every 60 seconds on the linked live pricing pages; confirm current cluster pricing before procuring.

Reserved vs On-Demand

On-Demand

Flexible, pay-by-the-hour, available immediately for single nodes. Best for experimentation, short runs, and bursty workloads. Priced at the top of the range, and constrained GPUs (H100/H200/B-series) can be unavailable at peak.

Reserved / Cluster

Two or more nodes on a 1-week to 12-month commitment, 16+ GPUs, typically 25–45% cheaper per GPU-hour, with guaranteed availability and a node-to-node fabric. Best for sustained training and production inference. Break-even versus on-demand is usually a few months of continuous use.

Interconnect & InfiniBand

Within a node, NVLink connects the 8 GPUs at terabytes per second. Between nodes, InfiniBand (commonly 400–800 Gb/s per GPU on modern fabrics) provides the low-latency RDMA bandwidth that makes multi-node training scale near-linearly.

For distributed training the interconnect is not optional — without it, scaling efficiency drops sharply past a single node because GPUs stall waiting on gradient communication. For single-node jobs or embarrassingly-parallel inference, standard Ethernet is often enough.

QuantaCloud

Need GPUs at scale?

Building out an inference fleet or training cluster? QuantaCloud brokers reserved capacity across multiple data center partners. 16+ GPUs, flexible terms, custom quote in 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

GPU Cluster FAQ

Citation

GPUPerHour: GPU Cluster Pricing Reference (September 2026)

Source: https://gpuperhour.com/gpu-cluster

6 GPU classes. Market ranges, manually verified.

28 live reserved cluster listings across 2 GPU products, refreshed hourly.

Last updated: September 2026

Per-GPU rates are tracked live across providers and update every 60 seconds. Cluster and reserved pricing varies by availability, region, term, and interconnect — confirm current rates before making procurement decisions.