GPU comparison

L40S vs MI300X

Specifications and current cloud pricing, side by side.

Ada LovelacevsCDNA 3Updated 17 days ago

The MI300X provides the stronger choice for the most common use case of LLM training. Its 192 GB memory, 5300 GB/s bandwidth, and 1307.4 TFLOPS dense FP16 rating together enable larger models and higher throughput than the L40S can sustain.

L40S from $0.80/GPU/hrMI300X from $2.99/GPU/hr

Right now, from live stock

  • Cheapest right now: L40S at $0.80/hr on Vast.ai

    Deploy
  • Most providers in stock: L40S (6)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: L40S ($0.0022 per TFLOPS-hour at 362.05 TFLOPS)

    Deploy

Specifications Compared

SpecL40SMI300X
TDP350W750W
VRAM48 GB192 GB
CUDA Cores18,176Not published
FP8 (dense)733 TFLOPS2,614.9 TFLOPS
Memory TypeGDDR6HBM3
ArchitectureAda LovelaceCDNA 3
FP16 (dense)362.05 TFLOPS1,307.4 TFLOPS
Form FactorsPCIeOAM
INT8 (dense)733 TOPS2,614.9 TOPS
InterconnectPCIe 4.0Infinity Fabric, PCIe 5.0
Tensor Cores568Not published
FP32 Performance91.6 TFLOPS163.4 TFLOPS
FP64 Performance1.4 TFLOPS81.7 TFLOPS
Memory Bandwidth864 GB/s5,300 GB/s
FP8 (with sparsity)1,466 TFLOPS5,229.8 TFLOPS
FP16 (with sparsity)733 TFLOPS2,614.9 TFLOPS
INT8 (with sparsity)1,466 TOPS5,229.8 TOPS

Performance Analysis

Dense FP16 throughput stands at 362.05 TFLOPS on the L40S and 1307.4 TFLOPS on the MI300X. The MI300X therefore supplies 3.6 times the dense FP16 operations per second. The same ratio appears when both GPUs run FP16 with sparsity at 733 TFLOPS versus 2614.9 TFLOPS. Training runs that fit within the 48 GB limit finish faster on the MI300X because its higher dense FP16 rating reduces iteration time. Memory bandwidth of 5300 GB/s on the MI300X versus 864 GB/s on the L40S permits batch sizes roughly six times larger before activation memory overflows. Inference workloads benefit similarly when models exceed 48 GB because the MI300X keeps entire layers resident. FP32 performance at 163.4 TFLOPS on the MI300X versus 91.6 TFLOPS on the L40S further widens the gap for mixed precision scientific kernels that rely on full precision accumulation.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

L40S

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiSlovenia, SI1$0.80—Deploy
Massed Computeus-central-21$0.97—Deploy
QuantaCloudus-midwest-14$1.09$4.36Deploy
RunPodglobal1$1.09—Deploy
LyceumEurope2$1.19$2.38Deploy

6 providers in stock, 13 offers (cheapest per provider shown). All L40S offers, price history and alerts

MI300X

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Hot AisleMichigan2$2.99$5.98Deploy

1 provider in stock, 2 offers (cheapest per provider shown). All MI300X offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when L40S drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.80/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the L40S

The L40S fits deployments that must stay inside a 350W power envelope because its TDP is less than half the MI300X rating. PCIe form factor cards also integrate directly into existing server chassis without OAM trays. Workloads that remain under 48 GB therefore complete at lower facility power cost on the L40S.

When to Choose the MI300X

The MI300X handles models that require 192 GB capacity because its memory is four times larger than the L40S. Bandwidth at 5300 GB/s supports batch sizes unattainable on the 864 GB/s L40S. Large scale training and inference therefore default to the MI300X when model size exceeds the smaller card limit.

Use Cases

LLM Training
MI300X

The MI300X supplies 192 GB memory and 1307.4 TFLOPS dense FP16 performance that exceed the L40S limits.

LLM Inference
MI300X

Larger batch sizes fit inside the 5300 GB/s bandwidth and 192 GB capacity of the MI300X.

Fine-tuning
MI300X

The 2614.9 TFLOPS FP16 with sparsity rating on the MI300X accelerates gradient updates beyond the L40S.

Stable Diffusion
L40S

The L40S at 350W TDP meets typical diffusion workloads that stay inside 48 GB without excess power draw.

Scientific Computing
MI300X

FP32 performance at 163.4 TFLOPS on the MI300X surpasses the 91.6 TFLOPS of the L40S for precision heavy codes.

Frequently Asked Questions

What are the dense FP16 ratings?▾

Dense FP16 reaches 362.05 TFLOPS on the L40S and 1307.4 TFLOPS on the MI300X.

Which GPU offers higher memory bandwidth?▾

The MI300X reaches 5300 GB/s while the L40S reaches 864 GB/s.

What TDP values are listed?▾

The L40S lists 350W TDP and the MI300X lists 750W TDP.

Do the cards share the same interconnect?▾

The L40S uses PCIe 4.0 while the MI300X uses Infinity Fabric together with PCIe 5.0.

Which architecture appears in each product?▾

The L40S uses Ada Lovelace while the MI300X uses CDNA 3.

Which is cheaper to rent, the L40S or the MI300X?▾

Cloud rental prices for both the L40S and MI300X vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the L40S have compared to the MI300X?▾

The L40S has 48 GB of GDDR6 memory. The MI300X has 192 GB of HBM3 memory.

Can I find L40S and MI300X GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the L40S and the MI300X?▾

The L40S uses the Ada Lovelace architecture (2023) while the MI300X uses CDNA 3 (2023). The MI300X delivers 3.6x the dense FP16 throughput (1,307.4 vs 362.05 TFLOPS, both without sparsity) and 6.1x the memory bandwidth of the L40S.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the L40S and the MI300X. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps