How much GPU memory 27 popular open-weight models need, and the smallest setup you can rent that holds each one. One row per model and precision. Every row assumes one request with an 8,192-token context; for anything else, put your own numbers into the LLM VRAM calculator, which uses the same arithmetic.
Last reviewed . Prices on this page are read live and are not part of the review date.
Two refinements, both read from the model's config. Layers that use a sliding or chunked attention window (Gemma 3, gpt-oss, Llama 4) stop growing their cache at the window size. DeepSeek's multi-head latent attention caches one compressed vector per token per layer instead of full keys and values. 1 GB is 1,000,000,000 bytes and a GPU sold as 80 GB is counted as 80 GB, which errs on the safe side.
This is a sizing floor for serving, not a benchmark. Long prompts, large batches and some serving engines need more; a total that lands within a few GB of a card's capacity is a sign to take the next size up.
| Model | Precision | Weights | KV cache (8k) | Total VRAM | Smallest setup | Cheapest setup right now |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | FP16as published | 16.1 GB | 1.1 GB | 18.8 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Llama 3.1 8B | FP8 | 8.0 GB | 1.1 GB | 10.0 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Llama 3.1 8B | INT4 | 4.0 GB | 1.1 GB | 5.6 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Llama 3.1 70B | FP16as published | 141 GB | 2.7 GB | 158 GB | 1x MI300X 192 GB$2.39 per hour on RunPod | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Llama 3.1 70B | FP8 | 70.6 GB | 2.7 GB | 80.6 GB | 1x H100 94 GB$3.11 per hour on Massed Compute | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Llama 3.1 70B | INT4 | 35.3 GB | 2.7 GB | 41.8 GB | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Llama 3.1 405B | FP16as published | 812 GB | 4.2 GB | 898 GB | 4x B300 262 GB$31.56 per hour on RunPod | 8x MI300X 192 GB$19.12 per hour on RunPod |
| Llama 3.1 405B | FP8 | 406 GB | 4.2 GB | 451 GB | 2x B300 262 GB$15.78 per hour on RunPod | 8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai |
| Llama 3.1 405B | INT4 | 203 GB | 4.2 GB | 228 GB | 1x B300 262 GB$7.89 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| Llama 3.3 70B | FP16as published | 141 GB | 2.7 GB | 158 GB | 1x MI300X 192 GB$2.39 per hour on RunPod | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Llama 3.3 70B | FP8 | 70.6 GB | 2.7 GB | 80.6 GB | 1x H100 94 GB$3.11 per hour on Massed Compute | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Llama 3.3 70B | INT4 | 35.3 GB | 2.7 GB | 41.8 GB | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Llama 4 Scout (17B active, 16 experts) | FP16as published | 217 GB | 1.6 GB | 241 GB | 1x B300 262 GB$7.89 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| Llama 4 Scout (17B active, 16 experts) | FP8 | 109 GB | 1.6 GB | 121 GB | 1x H200 141 GB$3.43 per hour on QuantaCloud | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Llama 4 Scout (17B active, 16 experts) | INT4 | 54.3 GB | 1.6 GB | 61.5 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Llama 4 Maverick (17B active, 128 experts) | FP16as published | 803 GB | 1.6 GB | 885 GB | 4x B300 262 GB$31.56 per hour on RunPod | 8x MI300X 192 GB$19.12 per hour on RunPod |
| Llama 4 Maverick (17B active, 128 experts) | FP8 | 402 GB | 1.6 GB | 444 GB | 2x B300 262 GB$15.78 per hour on RunPod | 8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai |
| Llama 4 Maverick (17B active, 128 experts) | INT4 | 201 GB | 1.6 GB | 223 GB | 1x B300 262 GB$7.89 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| Qwen2.5 7B | FP16as published | 15.2 GB | 0.5 GB | 17.3 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen2.5 7B | FP8 | 7.6 GB | 0.5 GB | 8.9 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen2.5 7B | INT4 | 3.8 GB | 0.5 GB | 4.7 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen2.5 32B | FP16as published | 65.5 GB | 2.1 GB | 74.4 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Qwen2.5 32B | FP8 | 32.8 GB | 2.1 GB | 38.4 GB | 1x A100 40 GB$1.27 per hour on VERDA | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Qwen2.5 32B | INT4 | 16.4 GB | 2.1 GB | 20.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen2.5 72B | FP16as published | 145 GB | 2.7 GB | 163 GB | 1x MI300X 192 GB$2.39 per hour on RunPod | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Qwen2.5 72B | FP8 | 72.7 GB | 2.7 GB | 82.9 GB | 1x H100 94 GB$3.11 per hour on Massed Compute | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Qwen2.5 72B | INT4 | 36.4 GB | 2.7 GB | 42.9 GB | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Qwen3 8B | FP16as published | 16.4 GB | 1.2 GB | 19.3 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen3 8B | FP8 | 8.2 GB | 1.2 GB | 10.3 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen3 8B | INT4 | 4.1 GB | 1.2 GB | 5.8 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen3 32B | FP16as published | 65.5 GB | 2.1 GB | 74.4 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Qwen3 32B | FP8 | 32.8 GB | 2.1 GB | 38.4 GB | 1x A100 40 GB$1.27 per hour on VERDA | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Qwen3 32B | INT4 | 16.4 GB | 2.1 GB | 20.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen3 30B-A3B | FP16as published | 61.1 GB | 0.8 GB | 68.1 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Qwen3 30B-A3B | FP8 | 30.5 GB | 0.8 GB | 34.5 GB | 1x A100 40 GB$1.27 per hour on VERDA | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Qwen3 30B-A3B | INT4 | 15.3 GB | 0.8 GB | 17.7 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Qwen3 235B-A22B | FP16as published | 470 GB | 1.6 GB | 519 GB | 2x B300 262 GB$15.78 per hour on RunPod | 8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai |
| Qwen3 235B-A22B | FP8 | 235 GB | 1.6 GB | 260 GB | 1x B300 262 GB$7.89 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| Qwen3 235B-A22B | INT4 | 118 GB | 1.6 GB | 131 GB | 1x H200 141 GB$3.43 per hour on QuantaCloud | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| DeepSeek V3 (671B) | FP16 | 1,342 GB | 0.6 GB | 1,477 GB | 8x MI300X 192 GB$19.12 per hour on RunPod | 8x MI300X 192 GB$19.12 per hour on RunPod |
| DeepSeek V3 (671B) | FP8as published | 671 GB | 0.6 GB | 739 GB | 4x MI300X 192 GB$9.56 per hour on RunPod | 8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai |
| DeepSeek V3 (671B) | INT4 | 336 GB | 0.6 GB | 370 GB | 2x MI300X 192 GB$4.78 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| DeepSeek R1 (671B) | FP16 | 1,342 GB | 0.6 GB | 1,477 GB | 8x MI300X 192 GB$19.12 per hour on RunPod | 8x MI300X 192 GB$19.12 per hour on RunPod |
| DeepSeek R1 (671B) | FP8as published | 671 GB | 0.6 GB | 739 GB | 4x MI300X 192 GB$9.56 per hour on RunPod | 8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai |
| DeepSeek R1 (671B) | INT4 | 336 GB | 0.6 GB | 370 GB | 2x MI300X 192 GB$4.78 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| DeepSeek R1 Distill Llama 8B | FP16as published | 16.1 GB | 1.1 GB | 18.8 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| DeepSeek R1 Distill Llama 8B | FP8 | 8.0 GB | 1.1 GB | 10.0 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| DeepSeek R1 Distill Llama 8B | INT4 | 4.0 GB | 1.1 GB | 5.6 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| DeepSeek R1 Distill Qwen 32B | FP16as published | 65.5 GB | 2.1 GB | 74.4 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| DeepSeek R1 Distill Qwen 32B | FP8 | 32.8 GB | 2.1 GB | 38.4 GB | 1x A100 40 GB$1.27 per hour on VERDA | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| DeepSeek R1 Distill Qwen 32B | INT4 | 16.4 GB | 2.1 GB | 20.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| DeepSeek R1 Distill Llama 70B | FP16as published | 141 GB | 2.7 GB | 158 GB | 1x MI300X 192 GB$2.39 per hour on RunPod | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| DeepSeek R1 Distill Llama 70B | FP8 | 70.6 GB | 2.7 GB | 80.6 GB | 1x H100 94 GB$3.11 per hour on Massed Compute | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| DeepSeek R1 Distill Llama 70B | INT4 | 35.3 GB | 2.7 GB | 41.8 GB | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Mistral 7B v0.3 | FP16as published | 14.5 GB | 1.1 GB | 17.1 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Mistral 7B v0.3 | FP8 | 7.3 GB | 1.1 GB | 9.2 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Mistral 7B v0.3 | INT4 | 3.6 GB | 1.1 GB | 5.2 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Mixtral 8x7B | FP16as published | 93.4 GB | 1.1 GB | 104 GB | 1x H200 141 GB$3.43 per hour on QuantaCloud | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Mixtral 8x7B | FP8 | 46.7 GB | 1.1 GB | 52.6 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Mixtral 8x7B | INT4 | 23.4 GB | 1.1 GB | 26.9 GB | 1x RTX 5090 32 GB$0.53 per hour on Vast.ai | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Mixtral 8x22B | FP16as published | 281 GB | 1.9 GB | 311 GB | 2x MI300X 192 GB$4.78 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
| Mixtral 8x22B | FP8 | 141 GB | 1.9 GB | 157 GB | 1x MI300X 192 GB$2.39 per hour on RunPod | 2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai |
| Mixtral 8x22B | INT4 | 70.3 GB | 1.9 GB | 79.4 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Mistral Small 24B (2501) | FP16as published | 47.1 GB | 1.3 GB | 53.3 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Mistral Small 24B (2501) | FP8 | 23.6 GB | 1.3 GB | 27.4 GB | 1x RTX 5090 32 GB$0.53 per hour on Vast.ai | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Mistral Small 24B (2501) | INT4 | 11.8 GB | 1.3 GB | 14.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Gemma 3 12B | FP16as published | 24.4 GB | 0.9 GB | 27.8 GB | 1x RTX 5090 32 GB$0.53 per hour on Vast.ai | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Gemma 3 12B | FP8 | 12.2 GB | 0.9 GB | 14.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Gemma 3 12B | INT4 | 6.1 GB | 0.9 GB | 7.7 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Gemma 3 27B | FP16as published | 54.9 GB | 1.1 GB | 61.6 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| Gemma 3 27B | FP8 | 27.4 GB | 1.1 GB | 31.4 GB | 1x RTX 5090 32 GB$0.53 per hour on Vast.ai | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Gemma 3 27B | INT4 | 13.7 GB | 1.1 GB | 16.3 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Phi-4 14B | FP16as published | 29.3 GB | 1.7 GB | 34.1 GB | 1x A100 40 GB$1.27 per hour on VERDA | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| Phi-4 14B | FP8 | 14.7 GB | 1.7 GB | 18.0 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| Phi-4 14B | INT4 | 7.3 GB | 1.7 GB | 9.9 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| gpt-oss 20B | MXFP4as published | 13.8 GB | 0.2 GB | 15.4 GB | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai | 1x RTX A5000 24 GB$0.23 per hour on Vast.ai |
| gpt-oss 20B | FP16 | 41.8 GB | 0.2 GB | 46.2 GB | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU | 1x RTX A6000 48 GB$0.44 per hour on LeaderGPU |
| gpt-oss 120B | MXFP4as published | 65.3 GB | 0.3 GB | 72.1 GB | 1x A100 80 GB$0.68 per hour on LeaderGPU | 1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai |
| gpt-oss 120B | FP16 | 234 GB | 0.3 GB | 257 GB | 1x B300 262 GB$7.89 per hour on RunPod | 4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai |
Smallest setup: fewest GPUs, then least combined VRAM, among GPUs with an offer right now. Cheapest setup: lowest hourly cost for the whole configuration, which is often more, smaller cards. Prices are the cheapest current on-demand price per GPU multiplied by the GPU count; a provider may not sell that exact count in one machine, so check the rent page. Prices observed .
One GPU, one request, 8,192 tokens of context. Each size lists what it adds over the size before it, so a bigger card also runs everything above its row.
| Single GPU | Models this size adds | GPUs with this much VRAM |
|---|---|---|
| 24 GB | Llama 3.1 8B FP16 (18.8 GB); Llama 3.1 8B FP8 (10.0 GB); Llama 3.1 8B INT4 (5.6 GB); Qwen2.5 7B FP16 (17.3 GB); Qwen2.5 7B FP8 (8.9 GB); Qwen2.5 7B INT4 (4.7 GB); Qwen2.5 32B INT4 (20.4 GB); Qwen3 8B FP16 (19.3 GB); Qwen3 8B FP8 (10.3 GB); Qwen3 8B INT4 (5.8 GB); Qwen3 32B INT4 (20.4 GB); Qwen3 30B-A3B INT4 (17.7 GB); DeepSeek R1 Distill Llama 8B FP16 (18.8 GB); DeepSeek R1 Distill Llama 8B FP8 (10.0 GB); DeepSeek R1 Distill Llama 8B INT4 (5.6 GB); DeepSeek R1 Distill Qwen 32B INT4 (20.4 GB); Mistral 7B v0.3 FP16 (17.1 GB); Mistral 7B v0.3 FP8 (9.2 GB); Mistral 7B v0.3 INT4 (5.2 GB); Mistral Small 24B (2501) INT4 (14.4 GB); Gemma 3 12B FP8 (14.4 GB); Gemma 3 12B INT4 (7.7 GB); Gemma 3 27B INT4 (16.3 GB); Phi-4 14B FP8 (18.0 GB); Phi-4 14B INT4 (9.9 GB); gpt-oss 20B MXFP4 (15.4 GB) | A10 from $0.37/hr, A30, L4 from $0.90/hr, Quadro P6000 from $1.10/hr, Quadro RTX 6000, RTX 3090 from $0.27/hr, RTX 4090 from $0.40/hr, RTX 4500 Ada, RTX A5000 from $0.23/hr |
| 48 GB | Llama 3.1 70B INT4 (41.8 GB); Llama 3.3 70B INT4 (41.8 GB); Qwen2.5 32B FP8 (38.4 GB); Qwen2.5 72B INT4 (42.9 GB); Qwen3 32B FP8 (38.4 GB); Qwen3 30B-A3B FP8 (34.5 GB); DeepSeek R1 Distill Qwen 32B FP8 (38.4 GB); DeepSeek R1 Distill Llama 70B INT4 (41.8 GB); Mixtral 8x7B INT4 (26.9 GB); Mistral Small 24B (2501) FP8 (27.4 GB); Gemma 3 12B FP16 (27.8 GB); Gemma 3 27B FP8 (31.4 GB); Phi-4 14B FP16 (34.1 GB); gpt-oss 20B FP16 (46.2 GB) | A40 from $0.49/hr, L40 from $0.86/hr, L40S from $0.80/hr, Quadro RTX 8000, RTX 5880 Ada, RTX 6000 Ada from $0.78/hr, RTX A6000 from $0.44/hr |
| 80 GB | Llama 4 Scout (17B active, 16 experts) INT4 (61.5 GB); Qwen2.5 32B FP16 (74.4 GB); Qwen3 32B FP16 (74.4 GB); Qwen3 30B-A3B FP16 (68.1 GB); DeepSeek R1 Distill Qwen 32B FP16 (74.4 GB); Mixtral 8x7B FP8 (52.6 GB); Mixtral 8x22B INT4 (79.4 GB); Mistral Small 24B (2501) FP16 (53.3 GB); Gemma 3 27B FP16 (61.6 GB); gpt-oss 120B MXFP4 (72.1 GB) | A100 from $0.68/hr, H100 from $2.50/hr |
| 141 GB | Llama 3.1 70B FP8 (80.6 GB); Llama 3.3 70B FP8 (80.6 GB); Llama 4 Scout (17B active, 16 experts) FP8 (121 GB); Qwen2.5 72B FP8 (82.9 GB); Qwen3 235B-A22B INT4 (131 GB); DeepSeek R1 Distill Llama 70B FP8 (80.6 GB); Mixtral 8x7B FP16 (104 GB) | H200 from $3.43/hr |
| 192 GB | Llama 3.1 70B FP16 (158 GB); Llama 3.3 70B FP16 (158 GB); Qwen2.5 72B FP16 (163 GB); DeepSeek R1 Distill Llama 70B FP16 (158 GB); Mixtral 8x22B FP8 (157 GB) | B200 from $3.75/hr, MI300X from $2.39/hr |
Full specs for each of these GPUs are in the GPU specs chart. FP8 rows run at full speed only on GPUs with FP8 hardware; the chart shows which publish an FP8 figure.
| Model | Parameters | Active | Layers | Hidden size | KV heads | Head dim | Context | Source |
|---|---|---|---|---|---|---|---|---|
| Llama 3.1 8BMeta | 8.03B | all (dense) | 32 | 4,096 | 8 | 128 | 131,072 | Model card, config |
| Llama 3.1 70BMeta | 70.55B | all (dense) | 80 | 8,192 | 8 | 128 | 131,072 | Model card, config |
| Llama 3.1 405BMeta | 405.85B | all (dense) | 126 | 16,384 | 8 | 128 | 131,072 | Model card, config |
| Llama 3.3 70BMeta | 70.55B | all (dense) | 80 | 8,192 | 8 | 128 | 131,072 | Model card, config |
| Llama 4 Scout (17B active, 16 experts)Meta | 108.64B | 17B | 48 | 5,120 | 8 | 128 | 10,485,760 | Model card, config |
| Llama 4 Maverick (17B active, 128 experts)Meta | 401.58B | 17B | 48 | 5,120 | 8 | 128 | 1,048,576 | Model card, config |
| Qwen2.5 7BAlibaba Qwen | 7.62B | all (dense) | 28 | 3,584 | 4 | 128 | 32,768 | Model card, config |
| Qwen2.5 32BAlibaba Qwen | 32.76B | all (dense) | 64 | 5,120 | 8 | 128 | 32,768 | Model card, config |
| Qwen2.5 72BAlibaba Qwen | 72.71B | all (dense) | 80 | 8,192 | 8 | 128 | 32,768 | Model card, config |
| Qwen3 8BAlibaba Qwen | 8.19B | all (dense) | 36 | 4,096 | 8 | 128 | 40,960 | Model card, config |
| Qwen3 32BAlibaba Qwen | 32.76B | all (dense) | 64 | 5,120 | 8 | 128 | 40,960 | Model card, config |
| Qwen3 30B-A3BAlibaba Qwen | 30.53B | 3.3B | 48 | 2,048 | 4 | 128 | 40,960 | Model card, config |
| Qwen3 235B-A22BAlibaba Qwen | 235.09B | 22B | 94 | 4,096 | 4 | 128 | 40,960 | Model card, config |
| DeepSeek V3 (671B)DeepSeek | 671B | 37B | 61 | 7,168 | MLA | 576 cached | 163,840 | Model card, config |
| DeepSeek R1 (671B)DeepSeek | 671B | 37B | 61 | 7,168 | MLA | 576 cached | 163,840 | Model card, config |
| DeepSeek R1 Distill Llama 8BDeepSeek | 8.03B | all (dense) | 32 | 4,096 | 8 | 128 | 131,072 | Model card, config |
| DeepSeek R1 Distill Qwen 32BDeepSeek | 32.76B | all (dense) | 64 | 5,120 | 8 | 128 | 131,072 | Model card, config |
| DeepSeek R1 Distill Llama 70BDeepSeek | 70.55B | all (dense) | 80 | 8,192 | 8 | 128 | 131,072 | Model card, config |
| Mistral 7B v0.3Mistral AI | 7.25B | all (dense) | 32 | 4,096 | 8 | 128 | 32,768 | Model card, config |
| Mixtral 8x7BMistral AI | 46.7B | 12.9B | 32 | 4,096 | 8 | 128 | 32,768 | Model card, config |
| Mixtral 8x22BMistral AI | 140.63B | 39B | 56 | 6,144 | 8 | 128 | 65,536 | Model card, config |
| Mistral Small 24B (2501)Mistral AI | 23.57B | all (dense) | 40 | 5,120 | 8 | 128 | 32,768 | Model card, config |
| Gemma 3 12BGoogle | 12.19B | all (dense) | 48 | 3,840 | 8 | 256 | 131,072 | Model card, config |
| Gemma 3 27BGoogle | 27.43B | all (dense) | 62 | 5,376 | 16 | 128 | 131,072 | Model card, config |
| Phi-4 14BMicrosoft | 14.66B | all (dense) | 40 | 5,120 | 10 | 128 | 16,384 | Model card, config |
| gpt-oss 20BOpenAI | 20.91B | 3.6B | 24 | 2,880 | 8 | 64 | 131,072 | Model card, config |
| gpt-oss 120BOpenAI | 116.83B | 5.1B | 36 | 2,880 | 8 | 64 | 131,072 | Model card, config |
Parameter totals are the safetensors parameter counts reported by each model's Hugging Face repository. Layers, hidden size, KV heads, head dimension and context length are the values in the repository's config.json, read on . Meta and Google gate their repositories, so for Llama and Gemma the config link goes to an ungated mirror of the same file. Context is the config's maximum position count; some model cards quote a different figure with or without rope scaling. Active parameter counts are the developer's published figures. DeepSeek V3 and R1 use the 671 billion main-model parameters the model card states.
GPU capacities come from GPUPerHour's datasheet-checked spec table, the same data as the GPU specs chart. Prices are read live when the page renders: the cheapest on-demand offer that is in stock from a secure provider, seen in the last 15 minutes, with MIG and fractional vGPU slices excluded. When either source cannot be reached the page says so and shows what it still can.
The table is free to reuse under CC BY 4.0 with a link to this page.
Llama 3.3 70B with an 8,192-token context needs about 158 GB at FP16, 80.6 GB at FP8, 41.8 GB at INT4, counting weights, KV cache and 10% overhead. The weights alone are 141 GB at FP16.
DeepSeek R1 has 671 billion parameters and is published in FP8, so its weights take 671 GB and the total with an 8,192-token context is about 739 GB. The smallest single-node setup that holds it is 4x MI300X 192 GB. Its multi-head latent attention keeps the KV cache small: 0.6 GB here.
Yes. The published MXFP4 checkpoint is 65.3 GB and the total with an 8,192-token context comes to about 72.1 GB.
At an 8,192-token context these fit in 24 GB: Llama 3.1 8B (FP16), Llama 3.1 8B (FP8), Qwen2.5 7B (FP16), Qwen2.5 7B (FP8), Qwen3 8B (FP16), Qwen3 8B (FP8), DeepSeek R1 Distill Llama 8B (FP16), DeepSeek R1 Distill Llama 8B (FP8), Mistral 7B v0.3 (FP16), Mistral 7B v0.3 (FP8), Gemma 3 12B (FP8), Phi-4 14B (FP8), gpt-oss 20B (MXFP4). With INT4 weights the list grows to models of roughly 30 billion parameters; see the table on this page.
Weights are parameters times bytes per parameter: 2 for FP16 or BF16, 1 for FP8, 0.5 for INT4. The KV cache is 2 x layers x KV heads x head dimension x 2 bytes x tokens. We add 10% on top for the CUDA context, activations and fragmentation. Every architecture number comes from the model's own config.json, linked in the facts table.
Every expert has to be in memory because the router can pick any of them for the next token. Active parameters decide how fast the model runs, not how much memory it needs. That is why the table sizes each model by its total parameter count.