L40 vs L40S

Ada LovelacevsAda LovelaceUpdated 6 days ago

The L40S serves as the stronger selection for the most common training and inference workloads. Its dense FP16 performance of 362.05 TFLOPS doubles the 181.05 TFLOPS figure of the L40 while memory capacity stays the same.

L40 from $0.86/GPU/hrL40S from $0.97/GPU/hr

Right now, from live stock

  • Cheapest right now: L40 at $0.86/hr on Massed Compute

    Deploy
  • Most providers in stock: L40S (4)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: L40S ($0.0027 per TFLOPS-hour at 362.05 TFLOPS)

    Deploy

Specifications Compared

SpecL40L40S
TDP300W350W
VRAM48 GB48 GB
CUDA Cores18,17618,176
FP8 (dense)362 TFLOPS733 TFLOPS
Memory TypeGDDR6GDDR6
ArchitectureAda LovelaceAda Lovelace
FP16 (dense)181.05 TFLOPS362.05 TFLOPS
Form FactorsPCIePCIe
INT8 (dense)362 TOPS733 TOPS
InterconnectPCIe 4.0PCIe 4.0
Tensor Cores568568
FP32 Performance90.5 TFLOPS91.6 TFLOPS
FP64 PerformanceNot published1.4 TFLOPS
Memory Bandwidth864 GB/s864 GB/s
FP8 (with sparsity)724 TFLOPS1,466 TFLOPS
FP16 (with sparsity)362.1 TFLOPS733 TFLOPS
INT8 (with sparsity)724 TOPS1,466 TOPS

Performance Analysis

The L40S supplies double the dense FP16 performance at 362.05 TFLOPS relative to 181.05 TFLOPS on the L40. This ratio permits the L40S to finish dense FP16 matrix calculations in half the duration during training workloads. Sparse FP16 performance maintains a comparable ratio at 733 TFLOPS on the L40S versus 362.1 TFLOPS on the L40. Memory bandwidth stays fixed at 864 GB/s on both units. Identical bandwidth keeps maximum batch sizes limited by the same transfer rate in memory intensive operations. FP32 performance exhibits little change at 91.6 TFLOPS versus 90.5 TFLOPS. The 350W TDP rating on the L40S accommodates its elevated tensor rates while the L40 functions at 300W.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

L40

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Massed Computeus-central-31$0.86Deploy
QuantaCloudus-midwest-22$0.94$1.88Deploy
HyperstackCANADA-11$1.00Deploy

3 providers in stock, 7 offers (cheapest per provider shown). All L40 offers, price history and alerts

L40S

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Massed Computeus-central-24$0.97$3.88Deploy
QuantaCloudus-midwest-18$1.09$8.72Deploy
RunPodglobal1$1.09Deploy
LyceumEurope1$1.19Deploy

4 providers in stock, 13 offers (cheapest per provider shown). All L40S offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when L40 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.86/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the L40

The L40 suits deployments where power limits constrain total system draw to 300W. Its FP32 performance of 90.5 TFLOPS remains nearly identical to the L40S value of 91.6 TFLOPS. This profile fits scientific computing tasks that emphasize FP32 throughput over tensor operations.

When to Choose the L40S

The L40S fits workloads that require the higher dense FP16 rate of 362.05 TFLOPS. Its sparse FP8 performance of 1,466 TFLOPS supports accelerated inference pipelines. The 350W TDP enables sustained operation at these tensor levels.

Use Cases

LLM Training
L40S

The L40S delivers dense FP16 performance of 362.05 TFLOPS compared with 181.05 TFLOPS on the L40.

LLM Inference
L40S

The L40S achieves dense FP8 performance of 733 TFLOPS against 362 TFLOPS on the L40.

Fine-tuning
L40S

The L40S provides sparse FP16 performance of 733 TFLOPS versus 362.1 TFLOPS on the L40.

Stable Diffusion
Either

Both units contain identical 48 GB GDDR6 memory and 864 GB/s bandwidth.

Scientific Computing
L40

The L40 operates at a 300W TDP with FP32 performance of 90.5 TFLOPS nearly matching the L40S value of 91.6 TFLOPS.

Frequently Asked Questions

What is the FP16 dense performance difference between the L40 and L40S?

The L40S records 362.05 TFLOPS while the L40 records 181.05 TFLOPS. This represents a two times ratio in dense FP16 throughput.

How does TDP differ between the L40 and L40S?

The L40S carries a 350W TDP rating. The L40 carries a 300W TDP rating.

Do the L40 and L40S share the same memory configuration?

Both models contain 48 GB GDDR6 memory with 864 GB/s bandwidth. Memory specifications remain identical.

What FP32 performance do the L40 and L40S deliver?

The L40S delivers 91.6 TFLOPS. The L40 delivers 90.5 TFLOPS.

How does sparse INT8 performance compare on the L40 versus L40S?

The L40S reaches 1,466 TOPS with sparsity. The L40 reaches 724 TOPS with sparsity.

Which model offers higher FP8 dense performance?

The L40S offers 733 TFLOPS in dense FP8. The L40 offers 362 TFLOPS in dense FP8.

Which is cheaper to rent, the L40 or the L40S?

Cloud rental prices for both the L40 and L40S vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the L40 have compared to the L40S?

The L40 has 48 GB of GDDR6 memory. The L40S has 48 GB of GDDR6 memory.

Can I find L40 and L40S GPUs available to rent right now?

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the L40 and the L40S?

The L40 uses the Ada Lovelace architecture (2023) while the L40S uses Ada Lovelace (2023). The L40S delivers 2.0x the dense FP16 throughput (362.05 vs 181.05 TFLOPS, both without sparsity) and the same memory bandwidth as the L40.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the L40 and the L40S. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps