NVIDIA alternatives fall into three access groups: hourly hardware rentals, chips rented from only one cloud, and hosted model APIs. Among these alternatives, AMD Instinct is the only established option rented by the hour across many independent clouds; start with the live MI300X rental table. The other main routes are chips rented from a single cloud, such as Google TPU and AWS Trainium, or API-first services such as Cerebras and Groq, with Intel Gaudi requiring a separate access check.
Choose the access model before comparing AI chips. If you need to run your own model and extensions, evaluate hardware and its software stack. If a hosted model already does the job, evaluate the endpoint. A fast chip is irrelevant if the service cannot run your workload.
NVIDIA competitors mapped to the access you need
This is an access map, not a stock list. The dated vendor descriptions identify routes to investigate; the live component below supplies tracked rental offers. "One cloud" describes the self-service route. "API only" describes the public inference route compared here, not every enterprise product a company sells.
| Accelerator | Maker | How you can use it | Where | Software stack | Dated fact from publisher |
|---|---|---|---|---|---|
| Instinct MI300X, MI325X, MI355X | AMD | Hourly rental across independent clouds | Vultr (all three), Hot Aisle (MI300X) and other rental providers | ROCm; PyTorch, vLLM, SGLang | Vultr's brief listed these families on 28 September 2026; Hot Aisle documented MI300X VMs on that date. |
| TPU | One cloud for self-service chip access | Google Cloud | JAX, PyTorch; XLA | Google's release notes date TPU7x general availability to 31 March 2026. | |
| Trainium | AWS | One cloud only | AWS EC2 | Neuron; TorchNeuron Native | AWS announced Trn3 UltraServer general availability on 2 December 2025. |
| Inferentia | AWS | One cloud only | AWS EC2 Inf instances | Neuron | AWS's instance history dates Inf2 introduction to 13 April 2023. |
| Gaudi2 and Gaudi 3 | Intel | Selected cloud access; not a broad hourly marketplace | Intel Tiber AI Cloud; IBM Cloud for Gaudi 3 | Confirm the offered framework image and SDK | Intel's presentation, accessed 28 September 2026, described Tiber access; IBM's documentation then labelled Gaudi 3 Select Availability. |
| Wafer-scale inference | Cerebras | API only for this comparison; dedicated deployments separately | Cerebras inference API | Hosted model API | Cerebras's public-model API listed gpt-oss-120b token prices on 28 September 2026. |
| LPU inference | Groq | API only for this comparison | GroqCloud | Hosted model API; LPU compiler underneath | GroqCloud's model documentation listed gpt-oss-120b token prices on 28 September 2026. |
The Gaudi row deliberately does not force selected cloud access into "many clouds" or "one cloud only." Neither label accurately describes Intel's and IBM's documented routes. For Gaudi, ask for the exact software image before treating an offer as a migration candidate.
Compare the tracked rentals live
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| MI300X | $3.39 | Hot Aisle | 1 |
| MI325X | none in stock | ||
| MI355X | none in stock | ||
| Intel Gaudi 2 | $0.91 | LeaderGPU | 1 |
| H100 | $2.50 | Hyperstack | 9 |
| H200 | $3.43 | QuantaCloud | 6 |
Read each row as a live per-GPU-hour comparison, then check the listing's full configuration and billing terms before comparing complete jobs.
Use the Gaudi2 rental page for its tracked offers. For a larger deployment, use the GPU cluster page to frame the whole-system requirement. Do not multiply a single-device result into a cluster purchase decision without testing the intended layout.
Set a baseline before changing hardware: the model, precision, input lengths, output lengths, concurrency and acceptable response time. Keep those fixed through the trial. Record completed useful work and total spend, including failed runs and migration time. A lower device rate is not enough if the new deployment misses your service target.
AMD: the first trial for memory-heavy inference
For an AMD AI chip evaluation, start with memory fit and the serving stack. Use the specifications below to narrow the family, then run your actual model. Do not select a larger device unless the extra capacity solves a measured constraint.
AMD's compatibility matrix dated 25 August 2026 lists PyTorch, vLLM and SGLang configurations. PyTorch's HIP documentation dated 11 August 2026 says its ROCm implementation reuses torch.cuda interfaces. That makes ordinary tensor code a reasonable starting point for a trial, not a promise that every extension will work.
AMD's HIPIFY documentation, accessed 28 September 2026, says it cannot translate a library without a HIP equivalent. SGLang's AMD documentation on that date also says its Marlin-dependent quantization paths do not work on AMD. Check the exact model format and kernels before moving the service.
My recommendation is to try AMD when memory limits your inference deployment and you can use a supported serving path. Budget a short correctness and load test. Read the Instinct family comparison for hardware selection and ROCm versus CUDA for migration work.
TPU and AWS: choose the cloud commitment first
Google TPU
Google's TPU documentation on 21 September 2026 recommends matrix-heavy models, large effective batches and long training runs. It identifies frequent branching, high-precision arithmetic and custom operations inside the training loop as poor fits. Treat that as the first filter for a TPU experiment.
Google's TPU7x documentation on the same date lists JAX and PyTorch support and explicitly excludes TensorFlow. Framework support must be checked for the specific generation, even when a broad product page lists your framework.
For a team already committed to Google Cloud, evaluate a representative training segment or serving workload before rewriting the deployment around TPUs. Use the TPU versus GPU guide for that comparison. Keep the decision tied to the code you plan to run, not the generation name.
AWS Trainium and Inferentia
AWS describes Inf2 as purpose-built for deep-learning inference on its product page, accessed 28 September 2026. For new PyTorch workloads, AWS's Neuron documentation on that date recommends TorchNeuron Native. It also says torch-neuronx is not included in Neuron 2.32.0 and later. An old migration recipe is therefore a poor basis for a new commitment.
Choose a Trainium or Inferentia trial when you intend to keep the workload inside AWS and can fund the software validation. Specify the instance family and SDK together. Do not assume that a successful build for one accelerator proves the same model will work on another.
The Trainium and Inferentia comparison covers that choice in depth. Use the AWS provider page for the site's tracked GPU comparison, without treating its GPU listings as Trainium quotes.
Intel Gaudi: evaluate an actual deployment offer
Intel Gaudi deserves a place on the shortlist when you have a concrete host and software image to evaluate. Keep Gaudi2 and Gaudi 3 separate. The live component tracks the former; it does not price the latter.
Intel's Gaudi presentation, accessed 28 September 2026, specifies 128 GB HBM for Gaudi 3. Capacity alone does not settle the decision. Ask the host to identify the supported model recipe, framework build, collective communication setup and upgrade policy. Run the same correctness and load checks you would use for AMD.
Do not start a Gaudi migration from a hardware headline. Start from an offer that names what you will actually run. If the provider cannot supply that, keep your existing deployment while the evaluation is unresolved.
Cerebras vs NVIDIA, and where Groq fits
For this decision, compare a hosted inference service with your existing serving deployment. These are token APIs, not GPU rentals tracked by this site. Test them when response time matters and their model choice fits your application.
Cerebras's public-model API on 28 September 2026 priced gpt-oss-120b at a derived $0.35 per million input tokens and $0.75 per million output tokens. Those figures multiply its published per-token prices by one million. GroqCloud's model documentation on 28 September 2026 listed the same model name at $0.15 per million input tokens and $0.60 per million output tokens.
Use those as dated billing references, not proof of equivalent quality or latency. Replay representative requests. Compare first-response time, completion time, output quality and behavior under concurrent load. Include your application's retries in the cost comparison.
Cerebras's 5 August 2025 announcement directs dedicated-capacity and custom-deployment inquiries to sales. Its dedicated-capacity route is therefore a separate evaluation from calling the public model API.
Groq also needs a corporate-context qualifier. Its 24 December 2025 announcement described a non-exclusive inference-technology licensing agreement with NVIDIA and said Groq would remain independent. Groq announced NVIDIA Cloud Partner membership on 12 August 2026. Treat Groq as an alternative service to evaluate, not a guarantee of permanent separation from NVIDIA technology. NVIDIA announced Groq 3 LPX inference racks as part of its Vera Rubin platform on 16 March 2026; NVIDIA, Groq and LPX covers what that means for serving models.
Smaller AI chip companies: separate trials from roadmaps
SambaNova's SambaStack page on 28 September 2026 described on-premises deployment and dedicated cloud hosting. Tenstorrent's documentation on that date advertised browser access through Cloud Console and named TT-Metalium as its low-level C++ SDK for Tensix kernels. These are distinct deployment paths, not interchangeable GPU offers.
d-Matrix announced a Demo Cloud for its Corsair inference hardware on 15 September 2026 and directed prospective users to a waitlist. Etched's website, accessed 28 September 2026, said it was validating its first rack-scale product with customers. Treat both as evaluation conversations, not assumptions in a production capacity plan.
Microsoft's 26 January 2026 Maia 200 announcement described an inference accelerator and an SDK preview. Meta's March 2026 announcement described MTIA 300 in production for ranking-and-recommendations training. Those announcements explain what the companies build for their workloads; they do not supply a public rental offer for yours.
ASIC vs GPU: specialization must fit the workload
The practical custom-accelerator trade is specialization versus flexibility. Google's TPU guidance gives a concrete example: favor matrix-dominated work, but reconsider workloads with frequent branching or custom operations in the training loop. Groq's 7 March 2025 architecture article describes a compiler that statically schedules data flow across chips, and an LPU that combines compute with on-chip SRAM.
The potential gain is a design organized around the work you need. The cost is fitting your model and execution path into that design. Neither source establishes a universal speedup over a GPU. Make compatibility the first gate and measured service performance the second.
Market share does not change that rule. ANALYST ESTIMATE: Kearney's 2 October 2025 article put NVIDIA's accelerator-market share at about 90% and forecast 70% by 2030. Its passage did not define a spending denominator, so those figures should not be read as measured GPU spending.
ANALYST ESTIMATE: Omdia's Alexander Harrowell, in an August 2026 forecast accessed 28 September 2026, expected custom AI-ASIC unit shipments to overtake GPUs in 2028. Omdia nevertheless expected GPU revenue to remain ahead through 2032. A unit forecast is not a forecast of your deployment economics.
Pick one trial with a clear exit condition
Try AMD first for memory-heavy inference with a supported stack. Try TPU or Trainium when the workload is committed to its cloud and worth validating there. Try Cerebras or Groq when response time is the constraint and the hosted model meets your needs. Pursue Gaudi when a specific deployment offer makes the software path concrete.
Stay on NVIDIA if your required extensions lack a supported alternative, your deadline leaves no room for validation, or the trial fails your quality and latency targets. Set the pass criteria before renting or integrating anything. Switch only when the complete workload passes and the savings or response-time improvement justify the migration.
Sources
- Vultr AMD solution brief, accessed 28 September 2026.
- Hot Aisle MI300X, accessed 28 September 2026.
- AMD ROCm compatibility matrix, 25 August 2026.
- PyTorch HIP semantics, 11 August 2026.
- AMD HIPIFY documentation, accessed 28 September 2026.
- SGLang AMD documentation, accessed 28 September 2026.
- Google TPU release notes, TPU7x entry 31 March 2026; accessed 21 September 2026.
- Google Cloud TPU product page, accessed 21 September 2026.
- Google TPU introduction and workload guidance, accessed 21 September 2026.
- Google TPU7x documentation, accessed 21 September 2026.
- AWS Trn3 announcement, 2 December 2025.
- AWS EC2 instance history, accessed 28 September 2026.
- AWS Inf2, accessed 28 September 2026.
- AWS Inferentia and Neuron, accessed 21 September 2026.
- AWS Neuron PyTorch integrations, accessed 28 September 2026.
- Intel Gaudi presentation, accessed 28 September 2026.
- IBM accelerated profiles, accessed 28 September 2026.
- Cerebras public-model prices, accessed 28 September 2026.
- Cerebras pricing field units, accessed 28 September 2026.
- Cerebras dedicated inference and gpt-oss announcement, 5 August 2025.
- GroqCloud models and pricing, accessed 28 September 2026.
- Groq LPU architecture, 7 March 2025.
- Groq and NVIDIA licensing agreement, 24 December 2025.
- Groq NVIDIA Cloud Partner announcement, 12 August 2026.
- SambaNova SambaStack, accessed 28 September 2026.
- Tenstorrent documentation, accessed 28 September 2026.
- d-Matrix Demo Cloud, 15 September 2026.
- Etched, accessed 28 September 2026.
- Microsoft Maia 200 announcement, 26 January 2026.
- Meta MTIA announcement, March 2026; accessed 28 September 2026.
- Kearney accelerator-market analysis, 2 October 2025.
- Omdia custom ASIC forecast, August 2026; accessed 28 September 2026.
- NVIDIA: Vera Rubin platform with Groq 3 LPX, 16 March 2026; accessed 1 October 2026.