GPU comparison

L40S vs RTX A6000

Specifications and current cloud pricing, side by side.

Ada LovelacevsAmpereUpdated 17 days ago

For the most common use case of LLM inference the L40S is the recommended choice. Its 362.05 TFLOPS dense FP16 performance exceeds the 154.8 TFLOPS dense FP16 performance of the RTX A6000 while memory capacity stays the same at 48 GB.

L40S from $0.80/GPU/hrRTX A6000 from $0.44/GPU/hr

Right now, from live stock

  • Cheapest right now: RTX A6000 at $0.44/hr on LeaderGPU

    Deploy
  • Most providers in stock: RTX A6000 (7)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: L40S ($0.0022 per TFLOPS-hour at 362.05 TFLOPS)

    Deploy

Specifications Compared

SpecL40SRTX A6000
TDP350W300W
VRAM48 GB48 GB
CUDA Cores18,17610,752
FP8 (dense)733 TFLOPSNot published
Memory TypeGDDR6GDDR6
ArchitectureAda LovelaceAmpere
FP16 (dense)362.05 TFLOPS154.8 TFLOPS
Form FactorsPCIePCIe
INT8 (dense)733 TOPS309.7 TOPS
InterconnectPCIe 4.0NVLink, PCIe 4.0
Tensor Cores568336
FP32 Performance91.6 TFLOPS38.7 TFLOPS
FP64 Performance1.4 TFLOPSNot published
Memory Bandwidth864 GB/s768 GB/s
FP8 (with sparsity)1,466 TFLOPSNot published
FP16 (with sparsity)733 TFLOPS309.6 TFLOPS
INT8 (with sparsity)1,466 TOPS619.4 TOPS

Performance Analysis

Dense FP16 performance reaches 362.05 TFLOPS on the L40S versus 154.8 TFLOPS on the RTX A6000. This gap means training and inference workloads that use dense FP16 operations complete more operations per second on the L40S. Sparse FP16 performance reaches 733 TFLOPS on the L40S versus 309.6 TFLOPS on the RTX A6000 so workloads that activate sparsity see a similar proportional advantage. FP32 performance stands at 91.6 TFLOPS for the L40S and 38.7 TFLOPS for the RTX A6000. Memory bandwidth of 864 GB/s on the L40S exceeds 768 GB/s on the RTX A6000 and therefore supports modestly larger batch sizes in memory bound phases.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

L40S

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiSlovenia, SI1$0.80—Deploy
Massed Computeus-central-31$0.97—Deploy
QuantaCloudus-midwest-14$1.09$4.36Deploy
RunPodglobal1$1.09—Deploy
LyceumEurope2$1.19$2.38Deploy

6 providers in stock, 14 offers (cheapest per provider shown). All L40S offers, price history and alerts

RTX A6000

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
LeaderGPUThe Netherlands8$0.44$3.54Deploy
QuantaCloudus-midwest-22$0.48$0.96Deploy
HyperstackCANADA-14$0.50$2.00Deploy
RunPodglobal1$0.53—Deploy
Massed Computeus-central-24$0.55$2.20Deploy

7 providers in stock, 58 offers (cheapest per provider shown). All RTX A6000 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when L40S drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.80/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the L40S

The L40S serves workloads that require 362.05 TFLOPS dense FP16 or 733 TFLOPS sparse FP16. Its 864 GB/s memory bandwidth further aids data movement in large batch inference. The 91.6 TFLOPS FP32 rating also exceeds the alternative by more than double.

When to Choose the RTX A6000

The RTX A6000 serves workloads that require the NVLink interconnect option. Its 300W TDP rating is lower than 350W and therefore reduces power draw in constrained environments. The 48 GB GDDR6 capacity remains identical to the L40S.

Use Cases

LLM Training
L40S

The L40S supplies 362.05 TFLOPS dense FP16 which exceeds the 154.8 TFLOPS dense FP16 of the RTX A6000.

LLM Inference
L40S

The L40S supplies 362.05 TFLOPS dense FP16 which exceeds the 154.8 TFLOPS dense FP16 of the RTX A6000.

Fine-tuning
L40S

The L40S supplies 362.05 TFLOPS dense FP16 which exceeds the 154.8 TFLOPS dense FP16 of the RTX A6000.

Stable Diffusion
L40S

The L40S supplies 362.05 TFLOPS dense FP16 which exceeds the 154.8 TFLOPS dense FP16 of the RTX A6000.

Scientific Computing
Either

Both GPUs provide 48 GB GDDR6 yet the L40S offers higher FP32 performance at 91.6 TFLOPS versus 38.7 TFLOPS.

Frequently Asked Questions

How does dense FP16 performance compare between the L40S and RTX A6000?▾

The L40S lists 362.05 TFLOPS dense FP16 while the RTX A6000 lists 154.8 TFLOPS dense FP16.

What is the FP32 performance of each GPU?▾

The L40S lists 91.6 TFLOPS FP32 and the RTX A6000 lists 38.7 TFLOPS FP32.

Which GPU has higher memory bandwidth?▾

The L40S provides 864 GB/s memory bandwidth while the RTX A6000 provides 768 GB/s memory bandwidth.

What interconnect options exist on the RTX A6000?▾

The RTX A6000 supports NVLink in addition to PCIe 4.0 while the L40S supports only PCIe 4.0.

How do the TDP ratings differ?▾

The L40S has a TDP of 350W and the RTX A6000 has a TDP of 300W.

Which is cheaper to rent, the L40S or the RTX A6000?▾

Cloud rental prices for both the L40S and RTX A6000 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the L40S have compared to the RTX A6000?▾

The L40S has 48 GB of GDDR6 memory. The RTX A6000 has 48 GB of GDDR6 memory.

Can I find L40S and RTX A6000 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the L40S and the RTX A6000?▾

The L40S uses the Ada Lovelace architecture (2023) while the RTX A6000 uses Ampere (2020). The L40S delivers 2.3x the dense FP16 throughput (362.05 vs 154.8 TFLOPS, both without sparsity) and 1.1x the memory bandwidth of the RTX A6000.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the L40S and the RTX A6000. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps