LLM GPU requirements

How much GPU memory 27 popular open-weight models need, and the smallest setup you can rent that holds each one. One row per model and precision. Every row assumes one request with an 8,192-token context; for anything else, put your own numbers into the LLM VRAM calculator, which uses the same arithmetic.

Last reviewed . Prices on this page are read live and are not part of the review date.

The formula and what it assumes

  • Weights = parameters x bytes per parameter. FP16 and BF16 use 2 bytes, FP8 uses 1, INT4 uses 0.5. For gpt-oss, which OpenAI publishes in MXFP4, we use the size of the published checkpoint.
  • KV cache = 2 x layers x KV heads x head dimension x bytes x tokens, with a 16-bit cache (2 bytes) and tokens = context length x concurrent sequences. Here that is 8,192 x 1.
  • Overhead = 10% of weights plus cache, for the CUDA context, activations and fragmentation.
  • GPU setup = the fewest GPUs of one type, out of 1, 2, 4 or 8, whose combined VRAM is at least the total. This assumes the model is split evenly across the GPUs with tensor parallelism. Only GPUs with 24 GB or more are considered.

Two refinements, both read from the model's config. Layers that use a sliding or chunked attention window (Gemma 3, gpt-oss, Llama 4) stop growing their cache at the window size. DeepSeek's multi-head latent attention caches one compressed vector per token per layer instead of full keys and values. 1 GB is 1,000,000,000 bytes and a GPU sold as 80 GB is counted as 80 GB, which errs on the safe side.

This is a sizing floor for serving, not a benchmark. Long prompts, large batches and some serving engines need more; a total that lands within a few GB of a card's capacity is a sign to take the next size up.

VRAM and minimum setup by model and precision

ModelPrecisionWeightsKV cache (8k)Total VRAMSmallest setupCheapest setup right now
Llama 3.1 8BFP16as published16.1 GB1.1 GB18.8 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Llama 3.1 8BFP88.0 GB1.1 GB10.0 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Llama 3.1 8BINT44.0 GB1.1 GB5.6 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Llama 3.1 70BFP16as published141 GB2.7 GB158 GB1x MI300X 192 GB$2.39 per hour on RunPod2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Llama 3.1 70BFP870.6 GB2.7 GB80.6 GB1x H100 94 GB$3.11 per hour on Massed Compute1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Llama 3.1 70BINT435.3 GB2.7 GB41.8 GB1x RTX A6000 48 GB$0.44 per hour on LeaderGPU1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Llama 3.1 405BFP16as published812 GB4.2 GB898 GB4x B300 262 GB$31.56 per hour on RunPod8x MI300X 192 GB$19.12 per hour on RunPod
Llama 3.1 405BFP8406 GB4.2 GB451 GB2x B300 262 GB$15.78 per hour on RunPod8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai
Llama 3.1 405BINT4203 GB4.2 GB228 GB1x B300 262 GB$7.89 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
Llama 3.3 70BFP16as published141 GB2.7 GB158 GB1x MI300X 192 GB$2.39 per hour on RunPod2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Llama 3.3 70BFP870.6 GB2.7 GB80.6 GB1x H100 94 GB$3.11 per hour on Massed Compute1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Llama 3.3 70BINT435.3 GB2.7 GB41.8 GB1x RTX A6000 48 GB$0.44 per hour on LeaderGPU1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Llama 4 Scout (17B active, 16 experts)FP16as published217 GB1.6 GB241 GB1x B300 262 GB$7.89 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
Llama 4 Scout (17B active, 16 experts)FP8109 GB1.6 GB121 GB1x H200 141 GB$3.43 per hour on QuantaCloud2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Llama 4 Scout (17B active, 16 experts)INT454.3 GB1.6 GB61.5 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Llama 4 Maverick (17B active, 128 experts)FP16as published803 GB1.6 GB885 GB4x B300 262 GB$31.56 per hour on RunPod8x MI300X 192 GB$19.12 per hour on RunPod
Llama 4 Maverick (17B active, 128 experts)FP8402 GB1.6 GB444 GB2x B300 262 GB$15.78 per hour on RunPod8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai
Llama 4 Maverick (17B active, 128 experts)INT4201 GB1.6 GB223 GB1x B300 262 GB$7.89 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
Qwen2.5 7BFP16as published15.2 GB0.5 GB17.3 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen2.5 7BFP87.6 GB0.5 GB8.9 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen2.5 7BINT43.8 GB0.5 GB4.7 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen2.5 32BFP16as published65.5 GB2.1 GB74.4 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Qwen2.5 32BFP832.8 GB2.1 GB38.4 GB1x A100 40 GB$1.27 per hour on VERDA1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Qwen2.5 32BINT416.4 GB2.1 GB20.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen2.5 72BFP16as published145 GB2.7 GB163 GB1x MI300X 192 GB$2.39 per hour on RunPod2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Qwen2.5 72BFP872.7 GB2.7 GB82.9 GB1x H100 94 GB$3.11 per hour on Massed Compute1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Qwen2.5 72BINT436.4 GB2.7 GB42.9 GB1x RTX A6000 48 GB$0.44 per hour on LeaderGPU1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Qwen3 8BFP16as published16.4 GB1.2 GB19.3 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen3 8BFP88.2 GB1.2 GB10.3 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen3 8BINT44.1 GB1.2 GB5.8 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen3 32BFP16as published65.5 GB2.1 GB74.4 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Qwen3 32BFP832.8 GB2.1 GB38.4 GB1x A100 40 GB$1.27 per hour on VERDA1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Qwen3 32BINT416.4 GB2.1 GB20.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen3 30B-A3BFP16as published61.1 GB0.8 GB68.1 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Qwen3 30B-A3BFP830.5 GB0.8 GB34.5 GB1x A100 40 GB$1.27 per hour on VERDA1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Qwen3 30B-A3BINT415.3 GB0.8 GB17.7 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Qwen3 235B-A22BFP16as published470 GB1.6 GB519 GB2x B300 262 GB$15.78 per hour on RunPod8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai
Qwen3 235B-A22BFP8235 GB1.6 GB260 GB1x B300 262 GB$7.89 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
Qwen3 235B-A22BINT4118 GB1.6 GB131 GB1x H200 141 GB$3.43 per hour on QuantaCloud2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
DeepSeek V3 (671B)FP161,342 GB0.6 GB1,477 GB8x MI300X 192 GB$19.12 per hour on RunPod8x MI300X 192 GB$19.12 per hour on RunPod
DeepSeek V3 (671B)FP8as published671 GB0.6 GB739 GB4x MI300X 192 GB$9.56 per hour on RunPod8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai
DeepSeek V3 (671B)INT4336 GB0.6 GB370 GB2x MI300X 192 GB$4.78 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
DeepSeek R1 (671B)FP161,342 GB0.6 GB1,477 GB8x MI300X 192 GB$19.12 per hour on RunPod8x MI300X 192 GB$19.12 per hour on RunPod
DeepSeek R1 (671B)FP8as published671 GB0.6 GB739 GB4x MI300X 192 GB$9.56 per hour on RunPod8x RTX PRO 6000 96 GB$5.28 per hour on Packet.ai
DeepSeek R1 (671B)INT4336 GB0.6 GB370 GB2x MI300X 192 GB$4.78 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
DeepSeek R1 Distill Llama 8BFP16as published16.1 GB1.1 GB18.8 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
DeepSeek R1 Distill Llama 8BFP88.0 GB1.1 GB10.0 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
DeepSeek R1 Distill Llama 8BINT44.0 GB1.1 GB5.6 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
DeepSeek R1 Distill Qwen 32BFP16as published65.5 GB2.1 GB74.4 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
DeepSeek R1 Distill Qwen 32BFP832.8 GB2.1 GB38.4 GB1x A100 40 GB$1.27 per hour on VERDA1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
DeepSeek R1 Distill Qwen 32BINT416.4 GB2.1 GB20.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
DeepSeek R1 Distill Llama 70BFP16as published141 GB2.7 GB158 GB1x MI300X 192 GB$2.39 per hour on RunPod2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
DeepSeek R1 Distill Llama 70BFP870.6 GB2.7 GB80.6 GB1x H100 94 GB$3.11 per hour on Massed Compute1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
DeepSeek R1 Distill Llama 70BINT435.3 GB2.7 GB41.8 GB1x RTX A6000 48 GB$0.44 per hour on LeaderGPU1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Mistral 7B v0.3FP16as published14.5 GB1.1 GB17.1 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Mistral 7B v0.3FP87.3 GB1.1 GB9.2 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Mistral 7B v0.3INT43.6 GB1.1 GB5.2 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Mixtral 8x7BFP16as published93.4 GB1.1 GB104 GB1x H200 141 GB$3.43 per hour on QuantaCloud2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Mixtral 8x7BFP846.7 GB1.1 GB52.6 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Mixtral 8x7BINT423.4 GB1.1 GB26.9 GB1x RTX 5090 32 GB$0.53 per hour on Vast.ai1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Mixtral 8x22BFP16as published281 GB1.9 GB311 GB2x MI300X 192 GB$4.78 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai
Mixtral 8x22BFP8141 GB1.9 GB157 GB1x MI300X 192 GB$2.39 per hour on RunPod2x RTX PRO 6000 96 GB$1.32 per hour on Packet.ai
Mixtral 8x22BINT470.3 GB1.9 GB79.4 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Mistral Small 24B (2501)FP16as published47.1 GB1.3 GB53.3 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Mistral Small 24B (2501)FP823.6 GB1.3 GB27.4 GB1x RTX 5090 32 GB$0.53 per hour on Vast.ai1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Mistral Small 24B (2501)INT411.8 GB1.3 GB14.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Gemma 3 12BFP16as published24.4 GB0.9 GB27.8 GB1x RTX 5090 32 GB$0.53 per hour on Vast.ai1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Gemma 3 12BFP812.2 GB0.9 GB14.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Gemma 3 12BINT46.1 GB0.9 GB7.7 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Gemma 3 27BFP16as published54.9 GB1.1 GB61.6 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
Gemma 3 27BFP827.4 GB1.1 GB31.4 GB1x RTX 5090 32 GB$0.53 per hour on Vast.ai1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Gemma 3 27BINT413.7 GB1.1 GB16.3 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Phi-4 14BFP16as published29.3 GB1.7 GB34.1 GB1x A100 40 GB$1.27 per hour on VERDA1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
Phi-4 14BFP814.7 GB1.7 GB18.0 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
Phi-4 14BINT47.3 GB1.7 GB9.9 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
gpt-oss 20BMXFP4as published13.8 GB0.2 GB15.4 GB1x RTX A5000 24 GB$0.23 per hour on Vast.ai1x RTX A5000 24 GB$0.23 per hour on Vast.ai
gpt-oss 20BFP1641.8 GB0.2 GB46.2 GB1x RTX A6000 48 GB$0.44 per hour on LeaderGPU1x RTX A6000 48 GB$0.44 per hour on LeaderGPU
gpt-oss 120BMXFP4as published65.3 GB0.3 GB72.1 GB1x A100 80 GB$0.68 per hour on LeaderGPU1x RTX PRO 6000 96 GB$0.66 per hour on Packet.ai
gpt-oss 120BFP16234 GB0.3 GB257 GB1x B300 262 GB$7.89 per hour on RunPod4x RTX PRO 6000 96 GB$2.64 per hour on Packet.ai

Smallest setup: fewest GPUs, then least combined VRAM, among GPUs with an offer right now. Cheapest setup: lowest hourly cost for the whole configuration, which is often more, smaller cards. Prices are the cheapest current on-demand price per GPU multiplied by the GPU count; a provider may not sell that exact count in one machine, so check the rent page. Prices observed .

What fits in 24, 48, 80, 141 and 192 GB

One GPU, one request, 8,192 tokens of context. Each size lists what it adds over the size before it, so a bigger card also runs everything above its row.

Single GPUModels this size addsGPUs with this much VRAM
24 GBLlama 3.1 8B FP16 (18.8 GB); Llama 3.1 8B FP8 (10.0 GB); Llama 3.1 8B INT4 (5.6 GB); Qwen2.5 7B FP16 (17.3 GB); Qwen2.5 7B FP8 (8.9 GB); Qwen2.5 7B INT4 (4.7 GB); Qwen2.5 32B INT4 (20.4 GB); Qwen3 8B FP16 (19.3 GB); Qwen3 8B FP8 (10.3 GB); Qwen3 8B INT4 (5.8 GB); Qwen3 32B INT4 (20.4 GB); Qwen3 30B-A3B INT4 (17.7 GB); DeepSeek R1 Distill Llama 8B FP16 (18.8 GB); DeepSeek R1 Distill Llama 8B FP8 (10.0 GB); DeepSeek R1 Distill Llama 8B INT4 (5.6 GB); DeepSeek R1 Distill Qwen 32B INT4 (20.4 GB); Mistral 7B v0.3 FP16 (17.1 GB); Mistral 7B v0.3 FP8 (9.2 GB); Mistral 7B v0.3 INT4 (5.2 GB); Mistral Small 24B (2501) INT4 (14.4 GB); Gemma 3 12B FP8 (14.4 GB); Gemma 3 12B INT4 (7.7 GB); Gemma 3 27B INT4 (16.3 GB); Phi-4 14B FP8 (18.0 GB); Phi-4 14B INT4 (9.9 GB); gpt-oss 20B MXFP4 (15.4 GB)A10 from $0.37/hr, A30, L4 from $0.90/hr, Quadro P6000 from $1.10/hr, Quadro RTX 6000, RTX 3090 from $0.27/hr, RTX 4090 from $0.40/hr, RTX 4500 Ada, RTX A5000 from $0.23/hr
48 GBLlama 3.1 70B INT4 (41.8 GB); Llama 3.3 70B INT4 (41.8 GB); Qwen2.5 32B FP8 (38.4 GB); Qwen2.5 72B INT4 (42.9 GB); Qwen3 32B FP8 (38.4 GB); Qwen3 30B-A3B FP8 (34.5 GB); DeepSeek R1 Distill Qwen 32B FP8 (38.4 GB); DeepSeek R1 Distill Llama 70B INT4 (41.8 GB); Mixtral 8x7B INT4 (26.9 GB); Mistral Small 24B (2501) FP8 (27.4 GB); Gemma 3 12B FP16 (27.8 GB); Gemma 3 27B FP8 (31.4 GB); Phi-4 14B FP16 (34.1 GB); gpt-oss 20B FP16 (46.2 GB)A40 from $0.49/hr, L40 from $0.86/hr, L40S from $0.80/hr, Quadro RTX 8000, RTX 5880 Ada, RTX 6000 Ada from $0.78/hr, RTX A6000 from $0.44/hr
80 GBLlama 4 Scout (17B active, 16 experts) INT4 (61.5 GB); Qwen2.5 32B FP16 (74.4 GB); Qwen3 32B FP16 (74.4 GB); Qwen3 30B-A3B FP16 (68.1 GB); DeepSeek R1 Distill Qwen 32B FP16 (74.4 GB); Mixtral 8x7B FP8 (52.6 GB); Mixtral 8x22B INT4 (79.4 GB); Mistral Small 24B (2501) FP16 (53.3 GB); Gemma 3 27B FP16 (61.6 GB); gpt-oss 120B MXFP4 (72.1 GB)A100 from $0.68/hr, H100 from $2.50/hr
141 GBLlama 3.1 70B FP8 (80.6 GB); Llama 3.3 70B FP8 (80.6 GB); Llama 4 Scout (17B active, 16 experts) FP8 (121 GB); Qwen2.5 72B FP8 (82.9 GB); Qwen3 235B-A22B INT4 (131 GB); DeepSeek R1 Distill Llama 70B FP8 (80.6 GB); Mixtral 8x7B FP16 (104 GB)H200 from $3.43/hr
192 GBLlama 3.1 70B FP16 (158 GB); Llama 3.3 70B FP16 (158 GB); Qwen2.5 72B FP16 (163 GB); DeepSeek R1 Distill Llama 70B FP16 (158 GB); Mixtral 8x22B FP8 (157 GB)B200 from $3.75/hr, MI300X from $2.39/hr

Full specs for each of these GPUs are in the GPU specs chart. FP8 rows run at full speed only on GPUs with FP8 hardware; the chart shows which publish an FP8 figure.

Model facts and where they come from

ModelParametersActiveLayersHidden sizeKV headsHead dimContextSource
Llama 3.1 8BMeta8.03Ball (dense)324,0968128131,072Model card, config
Llama 3.1 70BMeta70.55Ball (dense)808,1928128131,072Model card, config
Llama 3.1 405BMeta405.85Ball (dense)12616,3848128131,072Model card, config
Llama 3.3 70BMeta70.55Ball (dense)808,1928128131,072Model card, config
Llama 4 Scout (17B active, 16 experts)Meta108.64B17B485,120812810,485,760Model card, config
Llama 4 Maverick (17B active, 128 experts)Meta401.58B17B485,12081281,048,576Model card, config
Qwen2.5 7BAlibaba Qwen7.62Ball (dense)283,584412832,768Model card, config
Qwen2.5 32BAlibaba Qwen32.76Ball (dense)645,120812832,768Model card, config
Qwen2.5 72BAlibaba Qwen72.71Ball (dense)808,192812832,768Model card, config
Qwen3 8BAlibaba Qwen8.19Ball (dense)364,096812840,960Model card, config
Qwen3 32BAlibaba Qwen32.76Ball (dense)645,120812840,960Model card, config
Qwen3 30B-A3BAlibaba Qwen30.53B3.3B482,048412840,960Model card, config
Qwen3 235B-A22BAlibaba Qwen235.09B22B944,096412840,960Model card, config
DeepSeek V3 (671B)DeepSeek671B37B617,168MLA576 cached163,840Model card, config
DeepSeek R1 (671B)DeepSeek671B37B617,168MLA576 cached163,840Model card, config
DeepSeek R1 Distill Llama 8BDeepSeek8.03Ball (dense)324,0968128131,072Model card, config
DeepSeek R1 Distill Qwen 32BDeepSeek32.76Ball (dense)645,1208128131,072Model card, config
DeepSeek R1 Distill Llama 70BDeepSeek70.55Ball (dense)808,1928128131,072Model card, config
Mistral 7B v0.3Mistral AI7.25Ball (dense)324,096812832,768Model card, config
Mixtral 8x7BMistral AI46.7B12.9B324,096812832,768Model card, config
Mixtral 8x22BMistral AI140.63B39B566,144812865,536Model card, config
Mistral Small 24B (2501)Mistral AI23.57Ball (dense)405,120812832,768Model card, config
Gemma 3 12BGoogle12.19Ball (dense)483,8408256131,072Model card, config
Gemma 3 27BGoogle27.43Ball (dense)625,37616128131,072Model card, config
Phi-4 14BMicrosoft14.66Ball (dense)405,1201012816,384Model card, config
gpt-oss 20BOpenAI20.91B3.6B242,880864131,072Model card, config
gpt-oss 120BOpenAI116.83B5.1B362,880864131,072Model card, config
  • Llama 4 Scout (17B active, 16 experts): 36 of 48 layers use chunked attention over 8,192 tokens. The parameter total includes the vision encoder.
  • Llama 4 Maverick (17B active, 128 experts): 36 of 48 layers use chunked attention over 8,192 tokens. The parameter total includes the vision encoder.
  • DeepSeek V3 (671B): Multi-head latent attention caches 576 values per token per layer (kv_lora_rank 512 plus qk_rope_head_dim 64). The repository also holds a 14B multi-token-prediction module that is not counted here.
  • DeepSeek R1 (671B): Same architecture as DeepSeek V3: multi-head latent attention caching 576 values per token per layer.
  • Gemma 3 12B: Five of every six layers use a 1,024-token sliding window. The parameter total includes the vision encoder.
  • Gemma 3 27B: Five of every six layers use a 1,024-token sliding window. The parameter total includes the vision encoder.
  • gpt-oss 20B: Published with MXFP4 expert weights; the checkpoint size is the sum of the repository's safetensors files. Half the layers use a 128-token sliding window.
  • gpt-oss 120B: Published with MXFP4 expert weights; the checkpoint size is the sum of the repository's safetensors files. Half the layers use a 128-token sliding window.

Method and sources

Parameter totals are the safetensors parameter counts reported by each model's Hugging Face repository. Layers, hidden size, KV heads, head dimension and context length are the values in the repository's config.json, read on . Meta and Google gate their repositories, so for Llama and Gemma the config link goes to an ungated mirror of the same file. Context is the config's maximum position count; some model cards quote a different figure with or without rope scaling. Active parameter counts are the developer's published figures. DeepSeek V3 and R1 use the 671 billion main-model parameters the model card states.

GPU capacities come from GPUPerHour's datasheet-checked spec table, the same data as the GPU specs chart. Prices are read live when the page renders: the cheapest on-demand offer that is in stock from a secure provider, seen in the last 15 minutes, with MIG and fractional vGPU slices excluded. When either source cannot be reached the page says so and shows what it still can.

The table is free to reuse under CC BY 4.0 with a link to this page.

Questions

How much VRAM does a 70B model need?

Llama 3.3 70B with an 8,192-token context needs about 158 GB at FP16, 80.6 GB at FP8, 41.8 GB at INT4, counting weights, KV cache and 10% overhead. The weights alone are 141 GB at FP16.

What GPUs do I need to run DeepSeek R1?

DeepSeek R1 has 671 billion parameters and is published in FP8, so its weights take 671 GB and the total with an 8,192-token context is about 739 GB. The smallest single-node setup that holds it is 4x MI300X 192 GB. Its multi-head latent attention keeps the KV cache small: 0.6 GB here.

Does gpt-oss 120B fit on one 80 GB GPU?

Yes. The published MXFP4 checkpoint is 65.3 GB and the total with an 8,192-token context comes to about 72.1 GB.

What can I run on a 24 GB GPU without 4-bit quantisation?

At an 8,192-token context these fit in 24 GB: Llama 3.1 8B (FP16), Llama 3.1 8B (FP8), Qwen2.5 7B (FP16), Qwen2.5 7B (FP8), Qwen3 8B (FP16), Qwen3 8B (FP8), DeepSeek R1 Distill Llama 8B (FP16), DeepSeek R1 Distill Llama 8B (FP8), Mistral 7B v0.3 (FP16), Mistral 7B v0.3 (FP8), Gemma 3 12B (FP8), Phi-4 14B (FP8), gpt-oss 20B (MXFP4). With INT4 weights the list grows to models of roughly 30 billion parameters; see the table on this page.

How is the VRAM figure calculated?

Weights are parameters times bytes per parameter: 2 for FP16 or BF16, 1 for FP8, 0.5 for INT4. The KV cache is 2 x layers x KV heads x head dimension x 2 bytes x tokens. We add 10% on top for the CUDA context, activations and fragmentation. Every architecture number comes from the model's own config.json, linked in the facts table.

Why does a mixture-of-experts model need so much VRAM when only a few billion parameters are active?

Every expert has to be in memory because the router can pick any of them for the next token. Active parameters decide how fast the model runs, not how much memory it needs. That is why the table sizes each model by its total parameter count.