Specifications Compared
| Spec | L40S | MI300X |
|---|---|---|
| TDP | 350W | 750W |
| VRAM | 48 GB | 192 GB |
| CUDA Cores | 18,176 | Not published |
| FP8 (dense) | 733 TFLOPS | 2,614.9 TFLOPS |
| Memory Type | GDDR6 | HBM3 |
| Architecture | Ada Lovelace | CDNA 3 |
| FP16 (dense) | 362.05 TFLOPS | 1,307.4 TFLOPS |
| Form Factors | PCIe | OAM |
| INT8 (dense) | 733 TOPS | 2,614.9 TOPS |
| Interconnect | PCIe 4.0 | Infinity Fabric, PCIe 5.0 |
| Tensor Cores | 568 | Not published |
| FP32 Performance | 91.6 TFLOPS | 163.4 TFLOPS |
| FP64 Performance | 1.4 TFLOPS | 81.7 TFLOPS |
| Memory Bandwidth | 864 GB/s | 5,300 GB/s |
| FP8 (with sparsity) | 1,466 TFLOPS | 5,229.8 TFLOPS |
| FP16 (with sparsity) | 733 TFLOPS | 2,614.9 TFLOPS |
| INT8 (with sparsity) | 1,466 TOPS | 5,229.8 TOPS |
Performance Analysis
Dense FP16 throughput stands at 362.05 TFLOPS on the L40S and 1307.4 TFLOPS on the MI300X. The MI300X therefore supplies 3.6 times the dense FP16 operations per second. The same ratio appears when both GPUs run FP16 with sparsity at 733 TFLOPS versus 2614.9 TFLOPS. Training runs that fit within the 48 GB limit finish faster on the MI300X because its higher dense FP16 rating reduces iteration time. Memory bandwidth of 5300 GB/s on the MI300X versus 864 GB/s on the L40S permits batch sizes roughly six times larger before activation memory overflows. Inference workloads benefit similarly when models exceed 48 GB because the MI300X keeps entire layers resident. FP32 performance at 163.4 TFLOPS on the MI300X versus 91.6 TFLOPS on the L40S further widens the gap for mixed precision scientific kernels that rely on full precision accumulation.
Current On-Demand Offers
Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.
L40S
| Provider | Region | GPUs | Per GPU / hr | Instance / hr | Deploy |
|---|---|---|---|---|---|
| Vast.ai | Slovenia, SI | 1 | $0.80 | — | Deploy |
| Massed Compute | us-central-2 | 1 | $0.97 | — | Deploy |
| QuantaCloud | us-midwest-1 | 4 | $1.09 | $4.36 | Deploy |
| RunPod | global | 1 | $1.09 | — | Deploy |
| Lyceum | Europe | 2 | $1.19 | $2.38 | Deploy |
6 providers in stock, 13 offers (cheapest per provider shown). All L40S offers, price history and alerts
MI300X
1 provider in stock, 2 offers (cheapest per provider shown). All MI300X offers, price history and alerts
Notify me when L40S drops below a price
One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.80/GPU-hr.
QuantaCloud
Comparing providers? We broker across all of them.
Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.
When to Choose the L40S
The L40S fits deployments that must stay inside a 350W power envelope because its TDP is less than half the MI300X rating. PCIe form factor cards also integrate directly into existing server chassis without OAM trays. Workloads that remain under 48 GB therefore complete at lower facility power cost on the L40S.
When to Choose the MI300X
The MI300X handles models that require 192 GB capacity because its memory is four times larger than the L40S. Bandwidth at 5300 GB/s supports batch sizes unattainable on the 864 GB/s L40S. Large scale training and inference therefore default to the MI300X when model size exceeds the smaller card limit.
Use Cases
The MI300X supplies 192 GB memory and 1307.4 TFLOPS dense FP16 performance that exceed the L40S limits.
Larger batch sizes fit inside the 5300 GB/s bandwidth and 192 GB capacity of the MI300X.
The 2614.9 TFLOPS FP16 with sparsity rating on the MI300X accelerates gradient updates beyond the L40S.
The L40S at 350W TDP meets typical diffusion workloads that stay inside 48 GB without excess power draw.
FP32 performance at 163.4 TFLOPS on the MI300X surpasses the 91.6 TFLOPS of the L40S for precision heavy codes.
Frequently Asked Questions
What are the dense FP16 ratings?▾
Dense FP16 reaches 362.05 TFLOPS on the L40S and 1307.4 TFLOPS on the MI300X.
Which GPU offers higher memory bandwidth?▾
The MI300X reaches 5300 GB/s while the L40S reaches 864 GB/s.
What TDP values are listed?▾
The L40S lists 350W TDP and the MI300X lists 750W TDP.
Do the cards share the same interconnect?▾
The L40S uses PCIe 4.0 while the MI300X uses Infinity Fabric together with PCIe 5.0.
Which architecture appears in each product?▾
The L40S uses Ada Lovelace while the MI300X uses CDNA 3.
Which is cheaper to rent, the L40S or the MI300X?▾
Cloud rental prices for both the L40S and MI300X vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.
How much VRAM does the L40S have compared to the MI300X?▾
The L40S has 48 GB of GDDR6 memory. The MI300X has 192 GB of HBM3 memory.
Can I find L40S and MI300X GPUs available to rent right now?▾
Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.
What is the main difference between the L40S and the MI300X?▾
The L40S uses the Ada Lovelace architecture (2023) while the MI300X uses CDNA 3 (2023). The MI300X delivers 3.6x the dense FP16 throughput (1,307.4 vs 362.05 TFLOPS, both without sparsity) and 6.1x the memory bandwidth of the L40S.
Rent these GPUs
Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.
Related comparisons
How this page is made
- Specifications come from the NVIDIA, AMD and Intel datasheets for the L40S and the MI300X. Dense and sparse throughput are listed separately.
- Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
- The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
- Read how we collect the data, or report an error on this page.
Next steps
- Rent L40SEvery current offer by provider, daily price history and a price alert.
- Rent MI300XEvery current offer by provider, daily price history and a price alert.
- GPU price indexHow on-demand prices for the major GPUs have moved, updated daily.
- All GPUsEvery GPU we track, with current lows and provider counts.