AWS Trainium vs NVIDIA: Switch Only After a Pilot

Compare dated AWS Trainium and Inferentia list prices with live GPU rentals, check Neuron support, and decide whether migration pays for your workload.

By Faiz Ahmed•
•10 min read

Trainium is AWS's own chip for training AI models, and Inferentia is its chip for inference. On AWS's product pages on 28 September 2026, the single-chip trn1.2xlarge listed at USD 1.34 an hour, inf1.xlarge at USD 0.228 and inf2.xlarge at USD 0.76. They beat renting NVIDIA GPUs only when your model runs on AWS's Neuron software and the whole job, including the cost of porting it, comes out cheaper at the quality and speed you need.

Start with the software gate, then compare the bill. A cheaper chip-hour is useful only if it buys enough completed work. For a large, steady AWS workload, run a migration pilot. For a short job whose GPU implementation already works, rent the GPU unless the pilot shows a clear saving.

AWS Trainium and Inferentia list prices, with the billing unit exposed

AWS's Trn1, Inf1 and Inf2 product tables dated 28 September 2026 give the on-demand figures below in USD per whole instance-hour. AWS's product tables do not name a region for these prices, so confirm the rate in your region before committing.

The accelerator-hour column is derived: instance price divided by physical accelerator count, rounded to four decimal places. You still pay for the whole instance. AWS's same product tables supply the instance configurations; the Inf1 accelerator-memory totals are derived from its chip counts and Neuron's specification of 8 GiB per chip. Memory here means total accelerator memory, not host RAM.

InstanceAccelerator typeChipsTotal accelerator memoryOn-demand USD/instance-hourDerived USD/accelerator-hour
trn1.2xlargeTrainium132 GB1.341.3400
trn1.32xlargeTrainium16512 GB21.501.3438
trn1n.32xlargeTrainium16512 GB24.781.5488
inf1.xlargeInferentia18 GiB, derived0.2280.2280
inf1.2xlargeInferentia18 GiB, derived0.3620.3620
inf1.6xlargeInferentia432 GiB, derived1.1800.2950
inf1.24xlargeInferentia16128 GiB, derived4.7210.2951
inf2.xlargeInferentia2132 GB0.760.7600
inf2.8xlargeInferentia2132 GB1.971.9700
inf2.24xlargeInferentia26192 GB6.491.0817
inf2.48xlargeInferentia212384 GB12.981.0817

Do not pick an instance by the final column alone. The two single-chip Inf2 sizes have the same listed accelerator memory but different instance prices. Establish which complete instance sustains your service before comparing its cost. Likewise, unused chips still belong in the bill when your job cannot occupy the whole allocation.

This table has no Trn2 or Trn3 on-demand price. For Trainium2, AWS's Capacity Blocks pricing dated 28 September 2026 lists trn2.48xlarge in US East (Ohio) at USD 35.7608 per instance-hour. Dividing by its 16 chips gives a derived USD 2.23505 per accelerator-hour. That is a Capacity Block rate, not an ordinary on-demand rate. AWS's Trn2 product page on that date specifies 1.5 TB total accelerator memory for this instance.

Keep that reservation quote in a separate budget column. Do not fill a missing on-demand cell with it. AWS's EC2 pricing terms dated 28 September 2026 also state that prices generally exclude applicable taxes and duties. Use the AWS provider page when comparing the GPU alternative within AWS.

Trainium generations: memory and precision before peak compute

AWS's launch announcements and EC2 history give the launch dates below. AWS Neuron's hardware documentation dated 28 September 2026 supplies chip memory, bandwidth and the listed compute formats. Dates refer to the associated instance or UltraServer launch, not the first announcement of a chip design.

AcceleratorInstance launch or general availabilityMemory per chipMemory bandwidth per chipFormats in AWS's compute specifications
TrainiumTrn1: 10 October 202232 GiB HBM0.8 TB/secFP8, BF16, FP16, TF32
Trainium2Trn2: 3 December 202496 GiB2.9 TB/secFP8, BF16, FP16, TF32; separate sparse figures
Trainium3Trn3 UltraServer: 2 December 2025144 GiB4.9 TB/secMXFP8, MXFP4, BF16, FP16, TF32; separate sparse FP16/BF16/TF32 figures
InferentiaInf1: 3 December 20198 GiB DRAM50 GiB/secINT8, FP16, BF16
Inferentia2Inf2: 13 April 202332 GiB HBM820 GiB/secFP16, BF16, cFP8, TF32, INT8

AWS's own documents mix units: the Neuron architecture pages give memory in GiB and the product pages in GB or TB, and Trainium3 appears as 144 GiB in one place and 144 GB in another. The table keeps each figure as AWS printed it.

Use the table to shortlist a configuration, then test the model's actual memory demand. Do not infer model support from a precision label. Trainium 2 and Trainium3 hardware formats tell you what to investigate; they do not prove that your chosen model, quantization and serving path work together. The training versus inference guide helps separate the workload you need to size.

Neuron support is the migration gate

AWS's Neuron documentation dated 28 September 2026 identifies release 2.32.0, released on 17 August 2026. Its PyTorch guide recommends TorchNeuron Native for new workloads and documents eager execution, torch.compile and standard PyTorch distributed APIs. Start your evaluation there, with a pinned SDK version and a reproducible model configuration.

AWS's same PyTorch guide describes PyTorch NeuronX, torch-neuronx, as the XLA-based training and inference integration, but says it is not included in Neuron 2.32.0 and later. It also marks the Inf1 torch-neuron package as archived and no longer actively developed. An older tutorial is therefore insufficient evidence for a new deployment. Match its package names and versions to the stack you intend to operate.

For other frameworks, AWS's Inferentia page dated 21 September 2026 names PyTorch and TensorFlow integration through Neuron. AWS's JAX guide dated 28 September 2026 describes a PJRT plugin and labels JAX NeuronX beta, with some functionality potentially unsupported. Validate the exact instance and SDK combination; this is not a promise that every framework path covers every chip generation.

For serving, AWS's vLLM Neuron plugin documentation dated 28 September 2026 is labelled Beta and identifies Trn2 and Trn3 hardware. It documents an OpenAI-compatible API, continuous batching and EAGLE3 speculative decoding, plus GPT-OSS 20B and 120B deployment recipes. Do not extrapolate those Trn2/Trn3 recipes to Inferentia merely because both use Neuron.

Make the porting inventory concrete. Check model operators, custom kernels, distributed launch code, checkpoint loading, precision settings and serving behavior. AWS's Neuron overview dated 28 September 2026 documents NKI for custom kernels and a compiler that produces Neuron Executable File Format files, or NEFF. Budget for adapting kernels and compilation where your implementation needs them. Treat API compatibility as a starting point for validation, not proof of matching outputs or latency.

Compare live NVIDIA prices against completed work

The live table below supplies the NVIDIA rental side of the comparison.

GPUCheapest $/GPU-hrProviderProviders in stock
H100$2.50Hyperstack9
H200$3.43QuantaCloud6
B200$6.79RunPod3
L4$0.49RunPod2
A10$0.37LeaderGPU2
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Read it as per GPU-hour, against the AWS table's derived per accelerator-hour; equal hourly prices do not imply equal throughput or equal job cost. Open the H100 rental listings for the configuration behind a quote. For a distributed job, use the GPU cluster comparison to define the complete rental you need.

For training, compare total spend to reach the same target quality. Include compilation, warm-up, interrupted work and checkpoint recovery in your pilot budget. Keep the dataset, evaluation method and stopping criterion fixed. A faster run that misses the quality target is not a saving.

For inference, compare spend per accepted output under the same latency target. Use representative input lengths, output lengths and concurrency. Record idle time as well as busy time. If traffic is intermittent, include serverless GPU pricing in the shortlist before reserving a continuously billed instance.

Calculate migration payback separately: divide the one-time engineering cost by the measured operating saving over a fixed period. If the result exceeds the workload's expected lifetime, keep the working GPU deployment. Include ongoing maintenance in that calculation, especially if the team must support both implementations. The TPU versus GPU comparison provides another decision path if you are evaluating more than AWS accelerators.

Vendor claims and Rainier establish different things

VENDOR CLAIM: AWS's Trn1 product page dated 28 September 2026 advertises up to 50% cost-to-train savings against comparable EC2 instances. VENDOR CLAIM: Its Trn2 page on that date advertises 30 to 40% better price performance than EC2 P5e and P5en GPU instances. VENDOR CLAIM: Its Inf2 page advertises up to 40% better price performance than comparable EC2 instances on the same date. None establishes your saving against the live rentals above.

VENDOR CLAIM: AWS's 2 December 2025 Trn3 announcement advertises up to 4.4 times higher UltraServer performance and up to 4 times better performance per watt than Trn2 UltraServers. Those are system-level comparisons, not a measured price advantage for your model. Require a workload result before using either multiplier in a budget.

AWS reported on 3 November 2025 that Anthropic was already training Claude and running Claude inference on Project Rainier. Amazon's results on 5 February 2026 described Rainier as containing more than 500,000 Trainium2 chips. This is evidence of a large deployment. It is not a transferable benchmark for a smaller team or proof that the engineering effort will pay back on your workload.

Commit only after the workload passes

Choose Trainium for a large, sustained AWS training workload when the exact Neuron implementation passes quality and recovery checks and lowers total cost. Evaluate Trainium for inference too, but require the same serving checks you would apply to Inferentia. Choose Inferentia when its supported inference path meets your latency target and the complete instance bill wins.

Rent NVIDIA GPUs when your working implementation, short project lifetime or changing model requirements make migration hard to repay. For AWS Trainium vs NVIDIA, the decisive result is a reproducible workload bill. Commit to AWS accelerators only after the pilot beats the GPU baseline by enough to cover the port and its maintenance.

Sources

Frequently asked questions

What is the difference between Trainium and Inferentia?▾

AWS positions Trainium for training and Inferentia for inference. Trainium also runs inference: AWS reported Anthropic doing both on Project Rainier in November 2025.

Is AWS Trainium cheaper than NVIDIA GPUs?▾

Choose it only when your measured cost per completed training run or accepted inference output is lower after migration costs. AWS price-performance claims compare specific EC2 alternatives and do not establish savings against every GPU rental.

Can I use PyTorch on Trainium?▾

AWS's September 2026 documentation recommends TorchNeuron Native for new PyTorch workloads. Check the exact model, operators and distributed setup before committing.

Does vLLM support Trainium3?▾

AWS's September 2026 vLLM Neuron plugin documentation identifies Trn2 and Trn3 and labels the plugin Beta. It documents an OpenAI-compatible API and continuous batching.

Are Trainium2 Capacity Blocks on-demand prices?▾

No. Keep Capacity Block quotes separate from ordinary on-demand rates, and compare the region and reservation terms before budgeting.

Related Posts