GPU Cluster Pricing, 2026
A GPU cluster is more than one GPU node — typically 8 GPUs each, linked inside the node by NVLink — joined by a node-to-node fabric such as InfiniBand, RoCEv2 or high-speed Ethernet so they train and serve as one machine. This page covers what a cluster costs in 2026, how reserved cluster capacity compares to renting single nodes on-demand, and when each makes sense.
As a single-node baseline: an 8× H100 node on-demand runs roughly $15–28/hr, an 8× H200 node $24–36/hr, and an 8× A100 node $9–18/hr. Those are one node with no fabric between nodes. Multi-node reserved clusters on 3–12 month terms typically price 25–45% below on-demand per GPU-hour.
Live today: Lambda Labs reserved H100 SXM5 clusters from $5.54/GPU-hr and Lambda Labs reserved B200 SXM clusters from $8.87/GPU-hr, 16 to 1,536 GPUs on InfiniBand with a 1-week minimum. See the live cluster rates.
What Is a GPU Cluster?
A single node is one server, almost always 8 GPUs connected internally by NVLink. A cluster is two or more nodes connected by a node-to-node fabric — InfiniBand, RoCEv2 over Ethernet, or plain high-speed Ethernet — so dozens or thousands of GPUs cooperate on a single training or inference workload. NVLink never leaves the node; the fabric is what makes it a cluster.
Clusters are how teams pre-train and large-fine-tune models that don't fit on one node, and how high-throughput inference fleets are run with predictable capacity. The defining feature versus on-demand single GPUs is the interconnect: it keeps every GPU fed during the all-reduce communication that dominates distributed training.
Live GPU Cluster Pricing
Reserved multi-node cluster rates published by providers, verified against their pricing pages and refreshed hourly. Per-GPU-hour pricing steps down with cluster size; every listing names its node-to-node fabric and minimum term. Pick a size in the live table to compare it against on-demand nodes.
Lambda Labs · 16–512× H100 SXM5
8-GPU nodes · 80 GB VRAM per GPU · NVLink within each node · InfiniBand between nodes · 1-week minimum
| Cluster size | Per GPU-hour |
|---|---|
| 16–63 GPUs | $6.16 |
| 64–255 GPUs | $5.85 |
| 256+ GPUs | $5.54 |
Lambda Labs · 16–1536× B200 SXM
8-GPU nodes · 192 GB VRAM per GPU · NVLink within each node · InfiniBand between nodes · 1-week minimum
| Cluster size | Per GPU-hour |
|---|---|
| 16–63 GPUs | $9.86 |
| 64–255 GPUs | $9.36 |
| 256+ GPUs | $8.87 |
On-Demand Node Baseline
Not cluster prices: what a single 8-GPU node costs by the hour, with no fabric between nodes. This is the number a cluster quote is compared against. Per-GPU ranges are typical 2026 market rates; the cheapest live 8-GPU node comes from our tracker and refreshes hourly.
| GPU | On-Demand / GPU·hr | Cheapest 8-GPU Node Now |
|---|---|---|
| H100 SXM 80GB | $1.90–3.50 | $19.51/hrCoreWeave · United States · $2.44/GPU |
| H200 SXM 141GB | $3.00–4.50 | $20.64/hrCoreWeave · United States · $2.58/GPU |
| B200 SXM | $4.00–6.50 | $34.87/hrCoreWeave · United States · $4.36/GPU |
| GB300 NVL | $5.50–9.00 | Rack-scale (NVL72)typical range |
| A100 SXM 80GB | $1.10–2.20 | $9.20/hrDenvr · Virginia · $1.15/GPU |
| MI300X | $1.90–3.50 | $24.64/hrCirrascale · United States · $3.08/GPU |
Single-node on-demand, not cluster quotes. Per-GPU rates update every 60 seconds on the linked live pricing pages; confirm current cluster pricing before procuring.
Reserved vs On-Demand
On-Demand
Flexible, pay-by-the-hour, available immediately for single nodes. Best for experimentation, short runs, and bursty workloads. Priced at the top of the range, and constrained GPUs (H100/H200/B-series) can be unavailable at peak.
Reserved / Cluster
Two or more nodes on a 1-week to 12-month commitment, 16+ GPUs, typically 25–45% cheaper per GPU-hour, with guaranteed availability and a node-to-node fabric. Best for sustained training and production inference. Break-even versus on-demand is usually a few months of continuous use.
Interconnect & InfiniBand
Within a node, NVLink connects the 8 GPUs at terabytes per second. Between nodes, InfiniBand (commonly 400–800 Gb/s per GPU on modern fabrics) provides the low-latency RDMA bandwidth that makes multi-node training scale near-linearly.
For distributed training the interconnect is not optional — without it, scaling efficiency drops sharply past a single node because GPUs stall waiting on gradient communication. For single-node jobs or embarrassingly-parallel inference, standard Ethernet is often enough.
QuantaCloud
Need GPUs at scale?
Building out an inference fleet or training cluster? QuantaCloud brokers reserved capacity across multiple data center partners. 16+ GPUs, flexible terms, custom quote in 24 hours.
GPU Cluster FAQ
Citation
GPUPerHour: GPU Cluster Pricing Reference (September 2026)
Source: https://gpuperhour.com/gpu-cluster
6 GPU classes. Market ranges, manually verified.
28 live reserved cluster listings across 2 GPU products, refreshed hourly.
Last updated: September 2026
Per-GPU rates are tracked live across providers and update every 60 seconds. Cluster and reserved pricing varies by availability, region, term, and interconnect — confirm current rates before making procurement decisions.