AI inference is using a trained model to produce answers from new data; training is the process that sets the model's parameters in the first place, as IBM defines the two. On a GPU the difference is what runs out first: AWS describes training as typically compute-bound, and NVIDIA's inference guide splits inference into a compute-heavy step that reads the prompt (prefill) and a token-by-token step (decode) that is often limited by memory bandwidth. So size GPU memory first, then pay for dense compute for training and prefill, and for memory bandwidth for small-batch serving.
Training vs inference changes the job you rent
IBM describes training as a forward pass followed by a backward pass that improves parameters. Inference uses the forward pass. AWS's fine-tuning documentation, retrieved 28 September 2026, defines fine-tuning as further training that customizes an existing model by changing its weights. Treat full fine-tuning as a training-memory problem, not a cheaper-looking inference allocation.
The comparison below follows IBM's definitions, AWS's workload guidance and the ZeRO memory model. Billing describes rented compute, with Lambda's billing documentation, accessed 21 September 2026, as a concrete example: running instances remain billable while idle, and cluster reservations are priced per GPU-hour with contractual billing increments.
| Workload | What the GPU does | What limits speed | Typical GPU count or sizing basis | How long it runs | How rented compute is billed |
|---|---|---|---|---|---|
| Training | Forward and backward passes update parameters, per IBM | Typically compute throughput, per AWS; model states must fit in memory | Enough GPUs to hold the model states; ZeRO splits them across devices | A planned run with predictable resources, per AWS | Every allocated GPU for the run's billed duration; reservation terms apply |
| Fine-tuning | Further training changes existing weights, per AWS | The same as training when all weights change | Sized by the training method you choose | A training job | Allocated compute time under the rental contract |
| Inference | Forward passes produce predictions, per IBM | Prefill compute and often decode memory traffic, per NVIDIA | Size weights plus cache; NVIDIA's HGX architecture also supports large models across a node | Continuous or bursty real-time demand, per AWS | Instance lifetime for a persistent server; serverless billing rules vary |
AWS contrasts planned training runs with unpredictable inference traffic. For a rental decision, specify the model, precision, sequence length and concurrency before choosing a machine. For training, add the optimizer and gradient precision to that list.
LLM inference has two different bottlenecks
NVIDIA's November 2023 guide describes prefill as processing the known input tokens and preparing states for the first output token. Decode then generates subsequent tokens one at a time. A long prompt and a long answer therefore put different demands on the same server.
The Scaling Book's inference chapter, retrieved 28 September 2026, explains why: training and prefill reuse weights across many tokens. That gives the GPU more arithmetic for each byte read. NVIDIA's guide explains that moving weights, keys, values and activations can instead dominate decode latency. This is the hardware distinction behind AI inference vs training, not a rule that every inference operation is memory-bound.
Batching changes the balance. NVIDIA says a batch spreads weight reads across requests using the same model. The Scaling Book adds an important limit: sufficiently large batches can make decode's linear and feed-forward operations compute-bound, but each request keeps its own key-value, or KV, cache. Shared weights do not mean shared attention-cache traffic.
NVIDIA calls replacing finished sequences while others remain active continuous or in-flight batching. Use this distinction when evaluating a GPU for AI inference. For interactive serving, assess time to first token and the pace of later tokens separately. For an offline batch, prioritize completed work over the time one request spends waiting. Keep the prompt and output lengths fixed when comparing candidates.
Memory arithmetic before GPU shopping
The ZeRO paper, dated 13 May 2020, accounts for mixed-precision Adam using FP16 parameters and gradients plus FP32 master parameters, momentum and variance. Its model states consume 16 bytes per parameter. Activations, temporary buffers and fragmented memory are additional allocations.
Hugging Face's memory anatomy guide, retrieved 28 September 2026, uses FP32 gradients instead. Its example totals a derived 18 bytes per parameter: 6 for weights, 8 for Adam states and 4 for gradients. Neither total is a universal training-memory multiplier. Match the accounting to your implementation before spending money.
For inference, NVIDIA's 1 September 2026 sizing guidance gives two bytes per parameter for FP16/BF16 weights and one for FP8/INT8. Ideal packed 4-bit weights need a derived 0.5 byte per parameter, calculated as 4 divided by 8. NVIDIA's 24 June 2025 NVFP4 explanation includes scaling metadata, so that raw 4-bit calculation is only a lower bound.
Derived example: the same 7B parameters
This example uses an illustrative seven-billion-parameter model and decimal GB. It compares model states for training with raw weights for inference, not complete runtime allocations.
| Allocation | Derived arithmetic | Memory before other allocations |
|---|---|---|
| Mixed-precision Adam, ZeRO assumptions | 7 billion × 16 bytes | 112 GB |
| Mixed-precision Adam, Hugging Face assumptions | 7 billion × 18 bytes | 126 GB |
| FP16 inference weights | 7 billion × 2 bytes | 14 GB |
| FP8 inference weights | 7 billion × 1 byte | 7 GB |
| Ideal packed 4-bit inference weights | 7 billion × 0.5 byte | 3.5 GB |
Against the spec table's 80 GB H100 variant, both training totals exceed one GPU's capacity. ZeRO's full state partitioning ideally divides model-state storage by device count. A derived two-device split gives 56 GB or 63 GB per device for these respective totals, before activations and other allocations. That is a memory lower bound, not a validated training configuration.
The FP16 weights alone fit within the L4's 24 GB capacity. That does not establish a safe serving batch size. NVIDIA's November 2023 guide identifies KV-cache growth with sequence length and batch size. Leave space for the cache and runtime rather than treating all remaining memory as spare.
Use the LLM VRAM calculator to explore precision and context choices, then check the LLM GPU requirements reference. Record whether each estimate includes only weights, training states, or the whole runtime. Those are different budgets.
Read the hardware rows against the bottleneck
| Spec | B200 | H200 | H100 | L40S | L4 | RTX PRO 6000 |
|---|---|---|---|---|---|---|
| VRAM | 180 to 192 GB | 141 GB | 80 to 94 GB | 48 GB | 24 GB | 96 GB |
| Memory bandwidth | 8,000 GB/s | 4,800 GB/s | 3,350 GB/s | 864 GB/s | 300 GB/s | 1,792 GB/s |
| FP16 (dense) | 2,250 TFLOPS | 989 TFLOPS | 989 TFLOPS | 362.05 TFLOPS | 121 TFLOPS | 503.8 TFLOPS |
| FP16 (with sparsity) | 4,500 TFLOPS | 1,979 TFLOPS | 1,979 TFLOPS | 733 TFLOPS | 242 TFLOPS | 1,007.6 TFLOPS |
| FP8 (dense) | 4,500 TFLOPS | 1,979 TFLOPS | 1,979 TFLOPS | 733 TFLOPS | 242 TFLOPS | 1,007.6 TFLOPS |
| FP8 (with sparsity) | 9,000 TFLOPS | 3,958 TFLOPS | 3,958 TFLOPS | 1,466 TFLOPS | 485 TFLOPS | 2,015.2 TFLOPS |
| TDP | 1,000 W | 700 W | 700 W | 350 W | 72 W | 600 W |
Start with VRAM, then read bandwidth and dense compute separately. The site's verified family specifications give H100 and H200 the same dense FP16 and FP8 throughput, while H200 has more memory and bandwidth. That makes H200 a candidate when the memory budget or decode traffic is the constraint, not evidence of an automatic training speedup.
Use the dense FP16 or dense FP8 row for a dense workload. Sparse figures describe a different case. NVIDIA's 20 July 2021 sparsity guide requires a structured pattern of zeros and a pruning workflow. Do not use the sparse headline as the expected throughput of an unchanged model.
B200 leads these family rows in memory bandwidth and dense compute. L4, L40S and RTX PRO 6000 Blackwell offer other capacity and compute combinations worth checking against the actual job. These are family specifications, not a guarantee for every rented variant. Compare the exact offer before committing. The HBM vs GDDR guide explains the memory distinction; the precision guide helps choose which compute row to use.
Inference cost is an ongoing utilization decision
IBM Research's 5 October 2023 explanation describes training compute as an upfront investment and inference as ongoing work. For a time-based rental, budget a training run across every allocated GPU for its billed duration. Compare the total run, including the contract's billing increments, through the GPU cluster page.
The live table below shows current rental comparisons for the main candidates.
For an always-on inference server, use the monthly view below. It models continuous allocation rather than a forecast of request volume.
NVIDIA's September 2026 sizing guidance ties inference cost per token to utilization, memory, concurrency and latency targets. Compare the cost of completing your workload within those targets. A low rental total is not enough if the configuration misses the response-time requirement.
For bursty traffic, examine serverless GPU pricing. Runpod's documentation, accessed 21 September 2026, bills worker startup, execution and idle timeout. Serverless therefore needs its own billing calculation; do not assume that only useful token generation is charged.
Published cost trends measure different things
ANALYST ESTIMATE: Guido Appenzeller's a16z analysis, published 12 November 2024, estimated a tenfold annual inference-price decline at equivalent MMLU performance. It averaged input and output prices where they differed and sampled OpenAI, Anthropic and third-party-hosted Meta Llama. This is a defined API-price comparison, not a GPU rental forecast.
ANALYST ESTIMATE: Epoch AI's 22 September 2026 report estimated a 47% quarterly decline, equivalent to roughly thirteenfold annually, in the cost of fixed AI performance since 2023. Its benchmark set differs from a16z's MMLU comparison; do not average the two estimates into one forecast.
ANALYST ESTIMATE: Stanford's 2025 AI Index, retrieved 28 September 2026, reported a greater-than-280-fold inference-price reduction at GPT-3.5-level MMLU performance between November 2022 and October 2024. Keep those dates and the performance threshold attached to the claim. Use today's live rental comparisons for today's infrastructure budget.
The market split is still an estimate
ANALYST ESTIMATE: Deloitte's 18 November 2025 outlook projected inference at roughly two-thirds of AI compute in 2026. That is a compute-share forecast, not a measured spending split.
ANALYST ESTIMATE: McKinsey's 17 December 2025 model put 2025 inference data-center demand at 20.9 GW and training demand at 23.1 GW. It projected 93.3 GW for inference and 62.2 GW for training in 2030. Those power-demand estimates answer a different question from Deloitte's compute shares.
Neither estimate tells you which GPU your application needs. Use them as market context; size the rental from the workload.
Choose the GPU class from the work
For pre-training, shortlist H100, H200 and B200 cluster configurations. NVIDIA's HGX reference architecture, retrieved 28 September 2026, supports training and fine-tuning across these families. Favor dense compute after the training-state budget fits, then validate the intended distributed configuration before booking a long run.
For fine-tuning, calculate training memory first. Shortlist L40S for jobs that fit, and H100, H200 or B200 when the state budget requires more capacity or multiple GPUs. NVIDIA's L40S product page, retrieved 28 September 2026, positions it for both training and inference. Do not choose a full fine-tuning machine from the raw weight size alone.
For batch inference, start with L4 or L40S when the complete allocation fits. Compare H100, H200 and B200 if larger batches make compute or capacity the constraint. Use the same batch policy for every comparison.
For latency-sensitive serving, prioritize sufficient cache space and decode bandwidth. Shortlist H200 or B200 for demanding memory budgets, while keeping H100 in the live cost comparison. Use the best GPU for LLM guide to narrow the model pairing.
Rent the least expensive configuration that fits the complete memory budget and meets your measured latency or completion-time target. Pay for more dense compute when arithmetic limits the job; pay for more bandwidth when moving data limits token generation.
Sources
- IBM: AI inference, 13 January 2026.
- AWS: Fine-tuning foundation models, retrieved 28 September 2026.
- AWS: Inference compared with training, retrieved 28 September 2026.
- NVIDIA: Mastering LLM inference optimization, 17 November 2023.
- The Scaling Book: Inference, retrieved 28 September 2026.
- ZeRO paper, 13 May 2020.
- Hugging Face: Model memory anatomy, retrieved 28 September 2026.
- NVIDIA: GPU sizing guidance, 1 September 2026.
- NVIDIA: NVFP4 storage and precision, 24 June 2025.
- NVIDIA: Structured sparsity, 20 July 2021.
- IBM Research: Inference explained, 5 October 2023.
- Lambda: Cloud billing, accessed 21 September 2026.
- Runpod: Serverless billing, accessed 21 September 2026.
- a16z: LLMflation, 12 November 2024.
- Epoch AI: The plunging price of thought, 22 September 2026.
- Stanford: 2025 AI Index research and development, retrieved 28 September 2026.
- Deloitte: AI compute outlook, 18 November 2025.
- McKinsey: AI workloads and hyperscaler strategies, 17 December 2025.
- NVIDIA: HGX reference architecture, retrieved 28 September 2026.
- NVIDIA: L40S, retrieved 28 September 2026.