AI Inference vs Training: What Each Needs From a GPU

Understand AI inference, size training and serving memory, and choose a GPU using dense compute, memory bandwidth, workload limits and live rental prices.

By Faiz Ahmed•
•11 min read

AI inference is using a trained model to produce answers from new data; training is the process that sets the model's parameters in the first place, as IBM defines the two. On a GPU the difference is what runs out first: AWS describes training as typically compute-bound, and NVIDIA's inference guide splits inference into a compute-heavy step that reads the prompt (prefill) and a token-by-token step (decode) that is often limited by memory bandwidth. So size GPU memory first, then pay for dense compute for training and prefill, and for memory bandwidth for small-batch serving.

Training vs inference changes the job you rent

IBM describes training as a forward pass followed by a backward pass that improves parameters. Inference uses the forward pass. AWS's fine-tuning documentation, retrieved 28 September 2026, defines fine-tuning as further training that customizes an existing model by changing its weights. Treat full fine-tuning as a training-memory problem, not a cheaper-looking inference allocation.

The comparison below follows IBM's definitions, AWS's workload guidance and the ZeRO memory model. Billing describes rented compute, with Lambda's billing documentation, accessed 21 September 2026, as a concrete example: running instances remain billable while idle, and cluster reservations are priced per GPU-hour with contractual billing increments.

WorkloadWhat the GPU doesWhat limits speedTypical GPU count or sizing basisHow long it runsHow rented compute is billed
TrainingForward and backward passes update parameters, per IBMTypically compute throughput, per AWS; model states must fit in memoryEnough GPUs to hold the model states; ZeRO splits them across devicesA planned run with predictable resources, per AWSEvery allocated GPU for the run's billed duration; reservation terms apply
Fine-tuningFurther training changes existing weights, per AWSThe same as training when all weights changeSized by the training method you chooseA training jobAllocated compute time under the rental contract
InferenceForward passes produce predictions, per IBMPrefill compute and often decode memory traffic, per NVIDIASize weights plus cache; NVIDIA's HGX architecture also supports large models across a nodeContinuous or bursty real-time demand, per AWSInstance lifetime for a persistent server; serverless billing rules vary

AWS contrasts planned training runs with unpredictable inference traffic. For a rental decision, specify the model, precision, sequence length and concurrency before choosing a machine. For training, add the optimizer and gradient precision to that list.

LLM inference has two different bottlenecks

NVIDIA's November 2023 guide describes prefill as processing the known input tokens and preparing states for the first output token. Decode then generates subsequent tokens one at a time. A long prompt and a long answer therefore put different demands on the same server.

The Scaling Book's inference chapter, retrieved 28 September 2026, explains why: training and prefill reuse weights across many tokens. That gives the GPU more arithmetic for each byte read. NVIDIA's guide explains that moving weights, keys, values and activations can instead dominate decode latency. This is the hardware distinction behind AI inference vs training, not a rule that every inference operation is memory-bound.

Batching changes the balance. NVIDIA says a batch spreads weight reads across requests using the same model. The Scaling Book adds an important limit: sufficiently large batches can make decode's linear and feed-forward operations compute-bound, but each request keeps its own key-value, or KV, cache. Shared weights do not mean shared attention-cache traffic.

NVIDIA calls replacing finished sequences while others remain active continuous or in-flight batching. Use this distinction when evaluating a GPU for AI inference. For interactive serving, assess time to first token and the pace of later tokens separately. For an offline batch, prioritize completed work over the time one request spends waiting. Keep the prompt and output lengths fixed when comparing candidates.

Memory arithmetic before GPU shopping

The ZeRO paper, dated 13 May 2020, accounts for mixed-precision Adam using FP16 parameters and gradients plus FP32 master parameters, momentum and variance. Its model states consume 16 bytes per parameter. Activations, temporary buffers and fragmented memory are additional allocations.

Hugging Face's memory anatomy guide, retrieved 28 September 2026, uses FP32 gradients instead. Its example totals a derived 18 bytes per parameter: 6 for weights, 8 for Adam states and 4 for gradients. Neither total is a universal training-memory multiplier. Match the accounting to your implementation before spending money.

For inference, NVIDIA's 1 September 2026 sizing guidance gives two bytes per parameter for FP16/BF16 weights and one for FP8/INT8. Ideal packed 4-bit weights need a derived 0.5 byte per parameter, calculated as 4 divided by 8. NVIDIA's 24 June 2025 NVFP4 explanation includes scaling metadata, so that raw 4-bit calculation is only a lower bound.

Derived example: the same 7B parameters

This example uses an illustrative seven-billion-parameter model and decimal GB. It compares model states for training with raw weights for inference, not complete runtime allocations.

AllocationDerived arithmeticMemory before other allocations
Mixed-precision Adam, ZeRO assumptions7 billion × 16 bytes112 GB
Mixed-precision Adam, Hugging Face assumptions7 billion × 18 bytes126 GB
FP16 inference weights7 billion × 2 bytes14 GB
FP8 inference weights7 billion × 1 byte7 GB
Ideal packed 4-bit inference weights7 billion × 0.5 byte3.5 GB

Against the spec table's 80 GB H100 variant, both training totals exceed one GPU's capacity. ZeRO's full state partitioning ideally divides model-state storage by device count. A derived two-device split gives 56 GB or 63 GB per device for these respective totals, before activations and other allocations. That is a memory lower bound, not a validated training configuration.

The FP16 weights alone fit within the L4's 24 GB capacity. That does not establish a safe serving batch size. NVIDIA's November 2023 guide identifies KV-cache growth with sequence length and batch size. Leave space for the cache and runtime rather than treating all remaining memory as spare.

Use the LLM VRAM calculator to explore precision and context choices, then check the LLM GPU requirements reference. Record whether each estimate includes only weights, training states, or the whole runtime. Those are different budgets.

Read the hardware rows against the bottleneck

SpecB200H200H100L40SL4RTX PRO 6000
VRAM180 to 192 GB141 GB80 to 94 GB48 GB24 GB96 GB
Memory bandwidth8,000 GB/s4,800 GB/s3,350 GB/s864 GB/s300 GB/s1,792 GB/s
FP16 (dense)2,250 TFLOPS989 TFLOPS989 TFLOPS362.05 TFLOPS121 TFLOPS503.8 TFLOPS
FP16 (with sparsity)4,500 TFLOPS1,979 TFLOPS1,979 TFLOPS733 TFLOPS242 TFLOPS1,007.6 TFLOPS
FP8 (dense)4,500 TFLOPS1,979 TFLOPS1,979 TFLOPS733 TFLOPS242 TFLOPS1,007.6 TFLOPS
FP8 (with sparsity)9,000 TFLOPS3,958 TFLOPS3,958 TFLOPS1,466 TFLOPS485 TFLOPS2,015.2 TFLOPS
TDP1,000 W700 W700 W350 W72 W600 W
Figures from the vendor datasheets: B200, H200, H100, L40S, L4, RTX PRO 6000, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

Start with VRAM, then read bandwidth and dense compute separately. The site's verified family specifications give H100 and H200 the same dense FP16 and FP8 throughput, while H200 has more memory and bandwidth. That makes H200 a candidate when the memory budget or decode traffic is the constraint, not evidence of an automatic training speedup.

Use the dense FP16 or dense FP8 row for a dense workload. Sparse figures describe a different case. NVIDIA's 20 July 2021 sparsity guide requires a structured pattern of zeros and a pruning workflow. Do not use the sparse headline as the expected throughput of an unchanged model.

B200 leads these family rows in memory bandwidth and dense compute. L4, L40S and RTX PRO 6000 Blackwell offer other capacity and compute combinations worth checking against the actual job. These are family specifications, not a guarantee for every rented variant. Compare the exact offer before committing. The HBM vs GDDR guide explains the memory distinction; the precision guide helps choose which compute row to use.

Inference cost is an ongoing utilization decision

IBM Research's 5 October 2023 explanation describes training compute as an upfront investment and inference as ongoing work. For a time-based rental, budget a training run across every allocated GPU for its billed duration. Compare the total run, including the contract's billing increments, through the GPU cluster page.

The live table below shows current rental comparisons for the main candidates.

GPUCheapest $/GPU-hrProviderProviders in stock
B200$7.20VERDA1
H200$3.43QuantaCloud7
H100$2.59Vast.ai7
L40S$0.80Vast.ai4
L4$0.49RunPod3
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

For an always-on inference server, use the monthly view below. It models continuous allocation rather than a forecast of request volume.

GPU$/GPU-hrPer day (×24)Per month (×730)Provider
L4$0.49$11.76$358RunPod
L40S$0.80$19.20$584Vast.ai
H100$2.59$62.16$1,891Vast.ai
Cheapest in-stock on-demand price per GPU-hour, for one GPU running the whole time. Per day = hourly price × 24 hours. Per month = hourly price × 730 hours (365 days × 24 ÷ 12), rounded to the nearest dollar. Storage, data transfer and tax are not included. Latest stock observation: .

NVIDIA's September 2026 sizing guidance ties inference cost per token to utilization, memory, concurrency and latency targets. Compare the cost of completing your workload within those targets. A low rental total is not enough if the configuration misses the response-time requirement.

For bursty traffic, examine serverless GPU pricing. Runpod's documentation, accessed 21 September 2026, bills worker startup, execution and idle timeout. Serverless therefore needs its own billing calculation; do not assume that only useful token generation is charged.

ANALYST ESTIMATE: Guido Appenzeller's a16z analysis, published 12 November 2024, estimated a tenfold annual inference-price decline at equivalent MMLU performance. It averaged input and output prices where they differed and sampled OpenAI, Anthropic and third-party-hosted Meta Llama. This is a defined API-price comparison, not a GPU rental forecast.

ANALYST ESTIMATE: Epoch AI's 22 September 2026 report estimated a 47% quarterly decline, equivalent to roughly thirteenfold annually, in the cost of fixed AI performance since 2023. Its benchmark set differs from a16z's MMLU comparison; do not average the two estimates into one forecast.

ANALYST ESTIMATE: Stanford's 2025 AI Index, retrieved 28 September 2026, reported a greater-than-280-fold inference-price reduction at GPT-3.5-level MMLU performance between November 2022 and October 2024. Keep those dates and the performance threshold attached to the claim. Use today's live rental comparisons for today's infrastructure budget.

The market split is still an estimate

ANALYST ESTIMATE: Deloitte's 18 November 2025 outlook projected inference at roughly two-thirds of AI compute in 2026. That is a compute-share forecast, not a measured spending split.

ANALYST ESTIMATE: McKinsey's 17 December 2025 model put 2025 inference data-center demand at 20.9 GW and training demand at 23.1 GW. It projected 93.3 GW for inference and 62.2 GW for training in 2030. Those power-demand estimates answer a different question from Deloitte's compute shares.

Neither estimate tells you which GPU your application needs. Use them as market context; size the rental from the workload.

Choose the GPU class from the work

For pre-training, shortlist H100, H200 and B200 cluster configurations. NVIDIA's HGX reference architecture, retrieved 28 September 2026, supports training and fine-tuning across these families. Favor dense compute after the training-state budget fits, then validate the intended distributed configuration before booking a long run.

For fine-tuning, calculate training memory first. Shortlist L40S for jobs that fit, and H100, H200 or B200 when the state budget requires more capacity or multiple GPUs. NVIDIA's L40S product page, retrieved 28 September 2026, positions it for both training and inference. Do not choose a full fine-tuning machine from the raw weight size alone.

For batch inference, start with L4 or L40S when the complete allocation fits. Compare H100, H200 and B200 if larger batches make compute or capacity the constraint. Use the same batch policy for every comparison.

For latency-sensitive serving, prioritize sufficient cache space and decode bandwidth. Shortlist H200 or B200 for demanding memory budgets, while keeping H100 in the live cost comparison. Use the best GPU for LLM guide to narrow the model pairing.

Rent the least expensive configuration that fits the complete memory budget and meets your measured latency or completion-time target. Pay for more dense compute when arithmetic limits the job; pay for more bandwidth when moving data limits token generation.

Sources

Frequently asked questions

What is AI inference?▾

IBM defines inference as using a trained model to make predictions on new data. Training changes model parameters; inference uses them.

What is LLM inference?▾

LLM inference processes a prompt in prefill, then generates subsequent tokens during decode. NVIDIA describes decode as producing tokens one at a time.

Is inference always limited by memory bandwidth?▾

No. Small-batch decode often is, but prefill reuses weights across input tokens, and sufficiently large batches can make decode's feed-forward operations compute-bound.

How much memory does mixed-precision Adam training need?▾

ZeRO accounts for 16 bytes per parameter with FP16 gradients, before activations and buffers. Hugging Face's example uses FP32 gradients and totals 18 bytes per parameter before additional tensors.

Which GPU should I rent for AI inference?▾

Start with enough memory for weights, KV cache and runtime space. Then compare candidates against your latency and throughput targets using the live rental tables.

Related Posts