ROCm vs CUDA: What Runs on AMD GPUs and What Needs Porting

Check dated ROCm support for PyTorch, JAX and inference servers, identify CUDA porting work, and compare live GPU prices against measured workload costs.

By Faiz Ahmed•
•10 min read

ROCm is AMD's software stack for running compute work on its GPUs; CUDA is NVIDIA's. The current ROCm release is Core SDK 10.0.0, which AMD dates to 26 August 2026. The rule: standard PyTorch models and the major inference servers run on supported ROCm configurations with little or no code change, while custom CUDA kernels and CUDA-only libraries need porting, so audit those before you move.

Start with your dependencies, then compare prices. A model that imports successfully has passed an installation check. Your purchasing decision needs a completed training step or a serving run at the quality and latency you actually require. Treat those as separate gates, and set a migration budget before starting.

ROCm compatibility: pin the complete configuration

AMD's compatibility matrix dated 25 August 2026 lists validated combinations, not permission to mix arbitrary releases. AMD requires compatible firmware, kernel driver and user-space versions. Record those alongside the framework, model revision and container image for any trial you intend to repeat.

The dates below distinguish published documentation dates from documentation checked on 28 September 2026. A supported framework does not establish support for every extension inside your application.

Framework or libraryROCm status and boundaryPublisher and date
PyTorchAMD's ROCm 10.0.0 matrix lists 2.13.0, 2.12.0 and 2.11.0 configurations. PyTorch retains torch.cuda on HIP.AMD compatibility matrix, 25 August 2026; PyTorch HIP notes, 11 August 2026.
JAXAMD-supported ROCm plugin. AMD's matrix lists JAX 0.11.0, 0.10.2 and 0.10.0. Upstream installation instructions still describe a ROCm 7 plugin; verify the plugin-specific match before using SDK 10.0.0.AMD matrix, 25 August 2026; JAX installation docs, checked 28 September 2026.
TritonUpstream lists Linux and AMD GPUs with ROCm 6.2 or newer.Triton repository, checked 28 September 2026.
vLLMSupports MI300/gfx942 and MI350/gfx950. General minimum ROCm 6.3; MI350 needs 7.0 or newer. AMD's matrix lists vLLM 0.27.0.vLLM installation docs, checked 28 September 2026; AMD matrix, 25 August 2026.
SGLangAMD backend with Docker recommended; AMD's matrix lists 0.5.15. Marlin-dependent awq_marlin and gptq_marlin paths do not work on AMD.SGLang AMD docs, checked 28 September 2026; AMD matrix, 25 August 2026.
FlashAttentionROCm 6.0 or newer; Composable Kernel and Triton backends. FlashAttention-2 CK lists MI300X and MI355X, FP16/BF16, and forward/backward head dimensions up to 256.FlashAttention repository, checked 28 September 2026.
HIP and HIPIFYHIP is a portable GPU C++ dialect targeting AMD and NVIDIA. HIPIFY translates CUDA source into HIP C++; missing library equivalents remain a blocker.PyTorch HIP notes, 11 August 2026; AMD HIPIFY docs, checked 28 September 2026.
Windows and RadeonAMD's matrix includes Windows 11 25H2 for listed Radeon configurations. JAX lists Linux x86_64, experimental WSL2 and no native Windows GPU support.AMD matrix, 25 August 2026; JAX installation docs, checked 28 September 2026.

For a PyTorch ROCm trial, choose a listed combination first. For JAX, resolve the documented plugin mismatch before choosing the image. The upstream JAX instructions say the plugin targets a matching installed ROCm version; installing the Python extra does not install ROCm itself.

Do not interpret Windows support as identical coverage across frameworks. AMD's 25 August 2026 matrix attaches support to specific configurations. Use that exact combination as your starting point when evaluating an AMD GPU for AI on a workstation.

The porting bill starts outside ordinary tensor code

PyTorch's HIP notes dated 11 August 2026 explain why ordinary model code can need few or no source changes: ROCm deliberately reuses torch.cuda. Device strings remain cuda and cuda:0; replacing them with rocm is wrong. The same notes identify torch.version.hip as the way to distinguish a ROCm build. Audit the installed build before rewriting working application code.

For serving, start with the project's AMD installation path. vLLM's instructions expose /dev/kfd and /dev/dri to the container and point to MI300X tuning guidance. SGLang's AMD instructions recommend Docker. Follow those requirements when preparing your trial, then freeze the successful environment.

The expensive work begins where dependencies assume a particular backend. SGLang's AMD documentation says its AWQ implementation uses Triton dequantization rather than Marlin. A checkpoint format alone is therefore an incomplete compatibility check. Record the selected quantization backend, attention implementation and custom extensions, then test the exact combination you plan to serve.

AMD's HIPIFY documentation describes source translation, followed by code review, correctness testing, replacement of unsupported constructs and performance optimization. It cannot translate a library that has no HIP equivalent. Treat a successful conversion as the start of kernel validation. Budget separately for building the extension, matching outputs and recovering performance.

Also inspect installation side effects. FlashAttention's repository warns that installing bundled AITER can replace Triton; it documents AITER_USE_SYSTEM_TRITON=1 to retain a matched installation. That is a concrete reason to preserve the environment that passed validation.

Use a small acceptance run before a full migration. For training, require acceptable loss behavior, checkpoint recovery and repeatable step time. For inference, require output quality, latency under expected concurrency and sustained throughput. PyTorch's 11 August 2026 HIP notes document differences between MI300 and NVIDIA TF32 arithmetic. Choose tolerances deliberately; our precision guide helps frame that choice.

What changed after the December 2024 criticism

SemiAnalysis published its MI300X/H100/H200 training investigation on 22 December 2024 after a five-month evaluation. It found public AMD software materially less usable for those training tests than custom AMD development images. Its favorable MI300X results used a 21 December work-in-progress build on an unmerged branch, with BF16 Llama 3 8B and a four-layer Llama 3 70B proxy. Those results were not a stock-install comparison.

There is dated evidence of improvement. SemiAnalysis's 23 April 2025 follow-up reported substantial software progress and MI300 hardware entering PyTorch CI/CD. The upstream PyTorch workflow checked on 28 September 2026 confirms MI300 runners and regression testing every three hours. That confirms testing infrastructure; it does not settle compatibility for your private extensions.

For inference, SemiAnalysis's 16 February 2026 InferenceX v2 assessment found competitive MI355X versus B200 results using SGLang for DeepSeek-R1 FP8. The setup used disaggregated prefill with MoRI on AMD and Dynamo on NVIDIA. Results depended on the throughput-versus-interactivity operating point. The same assessment found weakness when MI355X combined FP4, disaggregated serving and wide expert parallelism. Do not turn either observation into a universal vendor ranking.

MLPerf adds evidence about specific submitted workloads. AMD's 4 June 2025 Training v5.0 report identifies Llama 2 70B LoRA fine-tuning on MI300X and MI325X, including MangoBoost's two-node/16-GPU and four-node/32-GPU MI300X submissions. In the MLCommons Training v6.0 supplemental discussion dated 16 June 2026, AMD identifies Flux.1 FP8 data-parallel training on 64 MI325X GPUs.

MLCommons published Inference v6.1 on 16 September 2026. AMD's report that day covers MI355X, MI350X and MI350P across language, reasoning, video and recommendation workloads. These are submission-coverage facts, not a claim that AMD wins every comparison. Use published setups to select a relevant trial, then measure your own acceptance criteria. The Instinct family comparison covers the hardware choice separately.

CUDA's head start and the limits of translation

NVIDIA's CUDA 1.0 announcement is dated 26 June 2007, giving that release more than nineteen years of history by this article's date. NVIDIA's archive checked on 28 September 2026 lists CUDA Toolkit 13.4.2, dated September 2026. NVIDIA's 18 March 2026 GTC coverage says CUDA served over 6 million developers: VENDOR CLAIM. That ecosystem figure cannot tell you how much work your dependency tree needs.

NVIDIA's CUDA SDK licence dated 26 January 2026, clause 1.2.8, restricts reverse engineering, decompiling or disassembling SDK-generated output for translation to non-NVIDIA platforms. That wording concerns generated output; it is not a blanket statement about all source-code porting. Tom's Hardware's corrected 4 March 2024 report says the restriction existed online since 2021 and entered installed documentation with CUDA 11.6.

ZLUDA describes its purpose as running unmodified CUDA applications on non-NVIDIA GPUs. Its maintainer's 29 June 2026 update announced version 6, reported continuing ML instruction and library fixes, and said commercial funding had ended, returning development to a weekend project. Those facts do not establish correctness for an arbitrary production workload. Do not make a migration budget depend on unverified drop-in compatibility.

Compare live prices after matching memory and output

Read AMD against NVIDIA at similar memory capacity using the spec table, rather than treating adjacent price rows as equal-memory pairs. Use the MI300X versus H100 comparison to inspect that particular tradeoff, and the MI300X rental page for its listings.

SpecMI300XH100MI325XH200MI355XB200
VRAM192 GB80 to 94 GB256 GB141 GB288 GB180 to 192 GB
Memory typeHBM3HBM3HBM3eHBM3eHBM3eHBM3e
Figures from the vendor datasheets: MI300X, H100, MI325X, H200, MI355X, B200, checked 13 Sep 2026. "Not published" means the vendor gives no figure.
GPUCheapest $/GPU-hrProviderProviders in stock
MI300Xnone in stock
H100$2.59QuantaCloud8
MI325Xnone in stock
H200$3.43QuantaCloud7
MI355Xnone in stock
B200$7.20VERDA1
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Compare the complete configuration needed to meet your target. Multiply the per-GPU price by the required GPU count, then include other billed resources. Keep model revision, precision, input and output lengths, concurrency, quality threshold and latency target fixed in the comparison. Our cloud GPU acceptance guide provides a place to continue the validation work.

Artificial Analysis's 8 June 2025 comparison used cost divided by measured throughput. AMD commissioned that report; it is not an unsponsored assessment. Applying that method, let R be whole-system hourly compute cost and Q be measured output tokens per second. The derived compute cost per million output tokens is R × 1,000,000 / (3,600 × Q). This excludes migration costs and assumes the measured throughput persists for the billed period.

The derived rule is equally useful without a currency figure: AMD's compute cost is lower when its AMD-to-NVIDIA cost ratio is smaller than its AMD-to-NVIDIA throughput ratio, at equal quality and latency. A lower listed rate alone does not answer that calculation.

For payback, estimate migration labor, validation compute and ongoing maintenance. Divide the one-time migration cost by positive recurring savings to estimate the recovery period. Choose the time unit consistently. Reject a port whose payback extends beyond the workload's expected life, and include repeat validation in the maintenance budget.

Port the supported workload; keep the costly dependency

Port when the supported ROCm build passes your acceptance run and measured savings repay migration within the workload's remaining life. Start with standard PyTorch or a documented serving configuration. Stay on CUDA when an essential extension lacks a workable replacement, performance recovery consumes the savings, or the delivery deadline leaves no room for validation.

If you can change the platform more broadly, the TPU versus GPU guide covers another decision path. For this one, require a working pinned environment and a positive payback calculation before committing the workload to AMD.

Sources

Frequently asked questions

What is ROCm?▾

ROCm is AMD's GPU computing software stack. AMD's documentation checked on 28 September 2026 identifies Core SDK 10.0.0, released 26 August 2026, as its latest documentation release.

Does PyTorch ROCm require changing torch.cuda calls?▾

PyTorch's HIP documentation says ROCm deliberately retains torch.cuda and cuda device strings. Ordinary tensor and model code can need few or no source changes, but custom extensions need a separate audit.

Can HIPIFY port every CUDA dependency?▾

No. AMD documents CUDA-to-HIP source translation, but HIPIFY cannot supply a missing HIP equivalent for a library. Budget for correctness testing and performance work after conversion.

Does ROCm support Windows and Radeon?▾

AMD's dated compatibility matrix includes selected Radeon configurations on Windows. Check the exact GPU, operating system and framework combination before choosing an image.

When is an AMD GPU cheaper for AI?▾

Choose AMD when measured cost per accepted workload is lower and the expected savings repay migration and ongoing maintenance. Compare complete systems at the same quality and latency targets.

Related Posts