The cheapest GPU per hour is rarely the best value. This page divides each GPU's cheapest live rental price by what you rent it for: gigabytes of memory, dense compute and memory bandwidth. The answer depends on which of those your job runs out of first, so all three are here, recomputed from live prices every hour.
Published 1 October 2026. Prices observed .
Fits the most model in memory for the money. Small cards often win; see the next list for large memory.
The same measure, limited to sizes that hold a large language model on one card.
Most arithmetic for the money: matters for training and for serving many requests at once.
Fastest memory for the money: sets the ceiling on tokens per second for a single request.
30 of 30 GPU families with a current offer. Select a column heading to sort.
| NVIDIA A10 | GDDR6 | $0.37 on LeaderGPU | 1.55¢ (24 GB) | 0.30¢ | Not published | 61.80¢ |
|---|---|---|---|---|---|---|
| NVIDIA A100 | HBM2e | $0.68 on LeaderGPU | 0.85¢ (80 GB) | 0.22¢ | Not published | 33.21¢ |
| NVIDIA A40 | GDDR6 | $0.49 on RunPod | 1.02¢ (48 GB) | 0.33¢ | Not published | 70.40¢ |
| NVIDIA B200 | HBM3e | $6.79 on RunPod | 3.54¢ (192 GB) | 0.30¢ | 0.15¢ | 84.88¢ |
| NVIDIA B300 | HBM3e | $7.89 on RunPod | 3.01¢ (262 GB) | 0.35¢ | 0.18¢ | 98.63¢ |
| Intel Gaudi 2 | HBM2e | $0.91 on LeaderGPU | 0.95¢ (96 GB) | Not published | Not published | 37.20¢ |
| NVIDIA GeForce GTX 1080 | GDDR5X | $0.60 on LeaderGPU | 5.45¢ (11 GB) | Not published | Not published | $1.88 |
| NVIDIA H100 | HBM3 | $2.50 on Hyperstack | 3.13¢ (80 GB) | 0.25¢ | 0.13¢ | 74.63¢ |
| NVIDIA H200 | HBM3e | $3.43 on QuantaCloud | 2.43¢ (141 GB) | 0.35¢ | 0.17¢ | 71.46¢ |
| NVIDIA L4 | GDDR6 | $0.49 on RunPod | 2.04¢ (24 GB) | 0.40¢ | 0.20¢ | $1.63 |
| NVIDIA L40 | GDDR6 | $0.79 on ThunderCompute | 1.65¢ (48 GB) | 0.44¢ | 0.22¢ | 91.44¢ |
| NVIDIA L40S | GDDR6 | $0.97 on Massed Compute | 2.02¢ (48 GB) | 0.27¢ | 0.13¢ | $1.12 |
| AMD Instinct MI300X | HBM3 | $3.39 on Hot Aisle | 1.77¢ (192 GB) | 0.26¢ | 0.13¢ | 63.96¢ |
| NVIDIA Quadro P4000 | GDDR5 | $0.51 on Paperspace | 6.38¢ (8 GB) | Not published | Not published | $2.10 |
| NVIDIA Quadro P5000 | GDDR5X | $0.78 on Paperspace | 4.88¢ (16 GB) | Not published | Not published | $2.71 |
| NVIDIA Quadro P6000 | GDDR5X | $1.10 on Paperspace | 4.58¢ (24 GB) | Not published | Not published | $2.55 |
| NVIDIA Quadro RTX 4000 | GDDR6 | $0.56 on Paperspace | 7.00¢ (8 GB) | 0.98¢ | Not published | $1.35 |
| NVIDIA Quadro RTX 5000 | GDDR6 | $0.82 on Paperspace | 5.13¢ (16 GB) | 0.92¢ | Not published | $1.83 |
| NVIDIA GeForce RTX 3060 | GDDR6 | $0.10 on Vast.ai | 0.85¢ (12 GB) | Not published | Not published | 28.42¢ |
| NVIDIA GeForce RTX 3090 | GDDR6X | $0.29 on LeaderGPU | 1.19¢ (24 GB) | 0.40¢ | Not published | 30.55¢ |
| NVIDIA RTX 4000 Ada Generation | GDDR6 | $0.28 on RunPod | 1.40¢ (20 GB) | Not published | Not published | 77.78¢ |
| NVIDIA GeForce RTX 4090 | GDDR6X | $0.74 on RunPod | 3.08¢ (24 GB) | 0.45¢ | 0.22¢ | 73.41¢ |
| NVIDIA GeForce RTX 5060 | GDDR7 | $0.18 on Vast.ai | 1.10¢ (16 GB) | Not published | Not published | 39.29¢ |
| NVIDIA GeForce RTX 5090 | GDDR7 | $0.53 on Vast.ai | 1.67¢ (32 GB) | 0.25¢ | 0.13¢ | 29.76¢ |
| NVIDIA RTX 6000 Ada Generation | GDDR6 | $0.78 on QuantaCloud | 1.62¢ (48 GB) | 0.21¢ | 0.11¢ | 80.86¢ |
| NVIDIA RTX A4000 | GDDR6 | $0.15 on Hyperstack | 0.94¢ (16 GB) | Not published | Not published | 33.48¢ |
| NVIDIA RTX A5000 | GDDR6 | $0.27 on RunPod | 1.13¢ (24 GB) | Not published | Not published | 35.16¢ |
| NVIDIA RTX A6000 | GDDR6 | $0.44 on LeaderGPU | 0.92¢ (48 GB) | 0.29¢ | Not published | 57.64¢ |
| NVIDIA RTX PRO 6000 Blackwell | GDDR7 | $1.56 on Vast.ai | 1.62¢ (96 GB) | 0.31¢ | 0.15¢ | 86.98¢ |
| NVIDIA Tesla V100 | HBM2 | $0.83 on Ori | 2.97¢ (32 GB) | 0.66¢ | Not published | 92.22¢ |
Costs are USD per hour for one unit: one GB of VRAM, one dense TFLOPS or one TB/s of bandwidth. Lower is better. "Not published" means the vendor gives no figure at that precision, so no value is computed. Prices observed , on-demand, in stock.
Start with memory. If the model, its KV cache and the runtime do not fit, nothing else matters. Use the per-GB list filtered to the size you need; the LLM VRAM calculator turns a model into a number.
Then ask what runs out first. Generating tokens for one user at a time is limited by memory bandwidth, so the per-TB/s list is the one to read for interactive serving. Training and batched serving are limited by arithmetic, so read the per-TFLOPS list, and the FP8 column if your software runs in FP8. The training versus inference guide and HBM versus GDDR explain why.
When the lists disagree, the job decides. A consumer card can top the per-GB and per-TFLOPS lists while lacking what a production job may need, such as NVLink between cards or a data-center memory system. The RTX PRO 6000 guide compares a professional and a consumer card built on the same chip.
Specs are per family. The spec table holds one row per GPU family, so the per-TFLOPS and per-TB/s figures divide the family's cheapest price by the family's published figure. Where a family has slower variants, such as the PCIe version of a GPU sold mainly as SXM, the cheapest offer may be the slower one, and the figure flatters it. The per-GB figure avoids this: it uses the price of the exact memory size. Check the GPU specs chart and the rent page before you book.
Dense figures only. Vendors headline throughput with structured sparsity, which is twice the dense number and applies only to pruned models. This page divides by the dense figure.
Which prices count. The price is the cheapest current on-demand offer per GPU-hour: in stock, from secure (non peer-to-peer) providers, seen in the last 15 minutes. MIG slices, fractional plans and laptop parts are left out, because a slice is not the whole GPU the specs describe. Storage, network transfer and tax are extra.
Specs come from GPUPerHour's spec table, served by the public /api/gpu-specs endpoint, each row checked against the vendor's datasheet. Prices come from live provider listings when the page renders. Price per GB is the cheapest price at each memory size divided by that size, taking the best size in the family. Price per TFLOPS is the family's cheapest price divided by its dense FP16 or FP8 figure. Price per TB/s divides by memory bandwidth in TB/s.
The figures are free to reuse under CC BY 4.0; please cite GPUPerHour and link to this page, with the observation time.
At the latest price observation (1 Oct 2026, 21:55 UTC), the NVIDIA A100: 0.85¢ per GB-hour (80 GB at $0.68 per GPU-hour), from LeaderGPU. Small consumer cards often lead this measure, so check the list filtered to the memory size your model needs.
At the latest price observation (1 Oct 2026, 21:55 UTC), the NVIDIA A100: 0.85¢ per GB-hour (80 GB at $0.68 per GPU-hour), from LeaderGPU.
At the latest price observation (1 Oct 2026, 21:55 UTC), measured on dense FP16 throughput, the NVIDIA RTX 6000 Ada Generation: 0.21¢ per TFLOPS-hour ($0.78 per GPU-hour), from QuantaCloud.
At the latest price observation (1 Oct 2026, 21:55 UTC), the NVIDIA GeForce RTX 3060: 28.42¢ per TB/s per hour ($0.10 per GPU-hour), from Vast.ai. Bandwidth sets the ceiling on tokens per second for a single request.
Vendors headline throughput with 2:4 structured sparsity, which is twice the dense figure and applies only to models pruned that way. An ordinary model runs at the dense rate, so the value figures here divide by the dense number.
The page reads live prices and recomputes at most once an hour. Prices move as providers change rates and stock sells out, so the rankings move too; every figure carries the time of the price observation it used.