Pick one of 27 open models or enter your own, set the precision, context and batch, and get the GPU memory it needs. Then see every setup of 1, 2, 4 or 8 GPUs that holds it, cheapest first, at today's rental prices.
Last reviewed . Prices on this page are read live and are not part of the review date.
VRAM needed
80.6 GB
| Setup | Combined VRAM | Cost per hour | FP8 figure published | Decode ceiling |
|---|---|---|---|---|
| 2x RTX A6000 48 GB | 96 GB | $0.89 on LeaderGPU | No | 22 tok/s |
| 1x Gaudi 2 96 GB | 96 GB | $0.91 on LeaderGPU | No | 35 tok/s |
| 2x A40 48 GB | 96 GB | $0.98 on RunPod | No | 20 tok/s |
| 4x RTX 3090 24 GB | 96 GB | $1.07 on Vast.ai | No | 53 tok/s |
| 4x RTX A5000 24 GB | 96 GB | $1.08 on RunPod | No | 44 tok/s |
| 4x L4 24 GB | 96 GB | $1.33 on Vast.ai | Yes | 17 tok/s |
| 2x A100 80 GB | 160 GB | $1.35 on LeaderGPU | No | 58 tok/s |
| 8x RTX 5060 16 GB | 128 GB | $1.41 on Vast.ai | No | 51 tok/s |
| 4x A10 24 GB | 96 GB | $1.48 on LeaderGPU | No | 34 tok/s |
| 2x RTX 6000 Ada 48 GB | 96 GB | $1.56 on QuantaCloud | Yes | 27 tok/s |
| 2x L40 48 GB | 96 GB | $1.72 on Massed Compute | Yes | 24 tok/s |
| 1x RTX PRO 6000 96 GB | 96 GB | $1.89 on VERDA | Yes | 25 tok/s |
| 8x RTX 2000 Ada 16 GB | 128 GB | $1.92 on RunPod | Yes | 25 tok/s |
| 2x L40S 48 GB | 96 GB | $1.94 on Massed Compute | Yes | 24 tok/s |
| 8x RTX A4000 16 GB | 128 GB | $2.00 on RunPod | No | 51 tok/s |
Llama 3.1 70B has 70.55 billion parameters, 80 layers, 8 KV heads and a head dimension of 128. At FP8 each parameter is one byte, so the weights are 70.6 GB. One 8,192-token sequence with a 16-bit cache adds 2.7 GB. Ten percent overhead on both is 7.3 GB, for a total of 80.6 GB. The cheapest setup that holds it right now is 2x RTX A6000 48 GB at $0.89 per hour on LeaderGPU.
| Setup | Combined VRAM | Cost per hour | Decode ceiling |
|---|---|---|---|
| 2x RTX A6000 48 GB | 96 GB | $0.89 on LeaderGPU | 22 tok/s |
| 1x Gaudi 2 96 GB | 96 GB | $0.91 on LeaderGPU | 35 tok/s |
| 2x A40 48 GB | 96 GB | $0.98 on RunPod | 20 tok/s |
| 4x RTX 3090 24 GB | 96 GB | $1.07 on Vast.ai | 53 tok/s |
| 4x RTX A5000 24 GB | 96 GB | $1.08 on RunPod | 44 tok/s |
| 4x L4 24 GB | 96 GB | $1.33 on Vast.ai | 17 tok/s |
Prices observed , USD, on-demand, in stock, per GPU times the GPU count.
Model presets and their sources are listed in full on LLM GPU requirements; every architecture number is from the model's Hugging Face config.json. GPU memory, bandwidth and the FP8 flag come from the datasheet-checked table behind the GPU specs chart. Prices are read live and never typed. The calculator runs in your browser; nothing you enter is sent anywhere.
Add three things. Weights: parameters times bytes per parameter (2 for FP16 or BF16, 1 for FP8, 0.5 for INT4). KV cache: 2 x layers x KV heads x head dimension x bytes per value x tokens, where tokens is context length times concurrent sequences. Overhead: 10% of the first two. For Llama 3.1 70B at FP8 with a 8,192-token context that is 70.6 GB + 2.7 GB + 7.3 GB = 80.6 GB.
It grows with every token you keep in context and every request you serve at once. Llama 3.1 70B stores 2 x 80 layers x 8 KV heads x 128 head dimension x 2 bytes per token, which comes to 2.7 GB for 8,192 tokens. Ten concurrent requests need ten times that. An FP8 cache halves it.
The calculator's rule: the frozen base weights at the precision you pick (INT4 is the QLoRA case), plus 16 bytes for every trainable parameter (16-bit weight and gradient, a 32-bit master copy and two Adam moments), plus checkpointed activations of batch x sequence length x hidden size x layers x 2 bytes, plus 10% overhead. It is a floor. Long sequences push real use higher because attention and MLP blocks are recomputed at full size.
Yes. Tensor parallelism splits the weights and the KV cache evenly across GPUs, so what matters is their combined VRAM. The calculator checks 1, 2, 4 and 8 GPUs of each type, which are the sizes providers rent. Splitting works best when the GPUs share a fast link such as NVLink; over plain PCIe it still fits but runs slower.
An upper bound on tokens per second for a single request. Generating a token reads every active weight once, so speed cannot exceed memory bandwidth divided by the size of the active weights. Real throughput is lower, and batching many requests raises total throughput well above it. It is there to compare GPUs, not to promise a number.
From live provider listings, read when the page is built and refreshed at least hourly: the cheapest on-demand price per GPU among offers that are in stock, from secure providers, seen in the last 15 minutes. The cost shown is that price times the number of GPUs. A provider may not sell that exact count in one machine, so follow the link to the rent page to check.