On-Demand, Spot or Reserved GPUs: How to Choose

How on-demand, spot, reserved, capacity block and marketplace GPU pricing work, what each one commits you to, and a rule for picking between them.

By Faiz Ahmed
13 min read

Use on-demand until a GPU is busy for most of the month. Use spot only for jobs that checkpoint and can be stopped with seconds of warning. Reserve only the capacity you would have run anyway. That is GPU as a service in three sentences. The rest of this page covers what each model bills you for, how much warning a spot GPU gives you, and what a reservation binds you to.

Pick a model by your situation

Your situationModel to useThe catch
Prototyping, debugging, bursty inference, anything shorter than a few weeksOn-demandHighest rate, and capacity can be sold out when you need it
Batch training, sweeps and offline inference that save checkpointsSpot, preemptible or interruptibleStopped with seconds of notice, and no service level agreement
A steady service or training programme that runs most hours for monthsReserved or committed useYou owe the money whether you use the GPU or not, often up front
A large run with a fixed start date, lasting days or weeksCapacity block or calendar reservationPaid in full up front, and AWS allows no cancellation
Price-sensitive single-GPU work that tolerates uneven hostsMarketplaceReliability, transfer charges and maximum rental length vary by listing
Multi-node training for a quarter or longerReserved cluster contractTake-or-pay terms, often multi-year

The gap between the models moves daily. The blocks below show the live spread between the cheapest on-demand, spot and reserved listing we track for the same GPU, starting with the H100.

Pricing typeCheapest listed $/GPU-hrvs on-demandProviderListingListings seen
On-demand$2.44baselineCoreWeave8 GPUs36 from 14 providers
Spot or preemptible$0.6374% lowerVast.ai1 GPU9 from 6 providers
Reserved$5.54127% higherLambda Labs512 GPUs, 1-week minimum term12 from 1 provider
Listed prices for H100: the cheapest catalogue listing of each pricing type seen in the last 24 hours, per GPU-hour, from secure (non-P2P) providers. Listed prices are not checked for stock, so they can differ from the in-stock on-demand price. Latest listing observation: .

The same spread for the A100, the older data centre card.

Pricing typeCheapest listed $/GPU-hrvs on-demandProviderListingListings seen
On-demand$0.68baselineLeaderGPU8 GPUs43 from 10 providers
Spot or preemptible$1.2484% higherMassed Compute8 GPUs1 from 1 provider
Reservedno reserved listing in the last 24 hours
Listed prices for A100: the cheapest catalogue listing of each pricing type seen in the last 24 hours, per GPU-hour, from secure (non-P2P) providers. Listed prices are not checked for stock, so they can differ from the in-stock on-demand price. Latest listing observation: .

And for the RTX 4090, a consumer card.

Pricing typeCheapest listed $/GPU-hrvs on-demandProviderListingListings seen
On-demand$0.51baselineVast.ai1 GPU9 from 3 providers
Spot or preemptible$0.4413% lowerVast.ai1 GPU4 from 2 providers
Reservedno reserved listing in the last 24 hours
Listed prices for RTX 4090: the cheapest catalogue listing of each pricing type seen in the last 24 hours, per GPU-hour, from secure (non-P2P) providers. Listed prices are not checked for stock, so they can differ from the in-stock on-demand price. Latest listing observation: .

These are listed prices, not stock-checked ones. A line reading "no listing in the last 24 hours" means no provider we track published that model for that GPU in that window, not that nobody sells it. Reserved pricing is often quote-only: Lambda's pricing page says "Contact us for reserved capacity at our lowest prices".

On-demand

You pay for the time an instance exists, you can end it whenever you like, and nobody can take the GPU mid-job. Runpod's wording is that resources "cannot be displaced by other users". The billing unit differs. Runpod Pods are "billed by the second for compute and storage". Vast.ai charges GPU compute "per second while your instance is running". Lambda bills "in one-minute increments", for as long as an instance runs, "regardless if they're actively being used". (All read on 21 September 2026.)

What goes wrong is rarely the hourly rate.

  • Stopped does not mean free. Runpod bills a stopped Pod's volume at twice the running storage rate. Vast.ai says "Storage charges continue even when instances are stopped." Lambda has no stop at all: instances "can only be launched, restarted, or terminated".
  • You may not get the GPU back. A stopped Runpod Pod stays tied to its machine. If someone else rents that GPU meanwhile, your Pod "cannot start with a GPU".
  • Capacity can be sold out. SemiAnalysis reported on 2 April 2026 that "On-Demand GPU rental capacity is sold out across all GPU types". On-demand is a price, not a guarantee of supply.
  • The rate itself moves. AWS announced on 5 June 2025 a reduction of "up to 45 percent" on its NVIDIA GPU instances. From 1 June 2025, On-Demand prices fell 33% on P4d and P4de (A100), 44% on P5 (H100) and 25% on P5en (H200), with Savings Plan reductions of 25% to 45% for purchases after 4 June. The on-demand side of any reserved comparison can drop by that much in one announcement.

Our daily GPU price index tracks how on-demand rates drift.

Spot, preemptible and interruptible

A spot GPU, a preemptible GPU and an interruptible GPU are the same trade: spare capacity sold cheaply, which the provider can take back. AWS and Azure call it Spot. Google Cloud calls it Spot VMs and used to call it preemptible. Vast.ai calls it interruptible.

What differs is the warning and what happens to the machine. These are the vendors' own documented terms, read on 21 September 2026.

ProviderDiscount the vendor statesWarningWhat happens to the instance
AWS Spot"up to a 90% discount" against On-DemandInterruption notice two minutes beforeStopped or terminated. No two minute warning if you use hibernation
Google Cloud Spot VMs"up to 91%" for many machine types, GPUs and TPUsShutdown period is "best effort and up to 30 seconds"No SLA. No maximum runtime (legacy preemptible VMs were capped at 24 hours)
Azure SpotNot captured in our researchEvicted "with 30-seconds notice"Deallocated by default, with disks still billed, or deleted. No SLA
Vast.ai interruptibleDocs say 50 to 80% savingsNone documented in the pages retrievedStopped, "killing running processes", and it "may wait long to resume"

Google Cloud's notice wording has changed. The current documentation gives a best effort shutdown period of up to 30 seconds, plus an optional 120 second preemption notice that is in Preview and defaults to 0 seconds. Google also says Spot prices "can change up to once every day".

Azure lets you set a maximum price. A value of -1 means you are never evicted for price.

Vast.ai's interruptible tier is an auction. Its docs say "You set a bid price" and the instance "Can be stopped by higher bids", while "On-demand instances always have highest priority". An outbid instance is paused, not deleted, so its storage keeps billing.

Runpod is a case we could not settle. A 5 second SIGTERM warning before SIGKILL is attributed to Runpod's Pod management docs, but we have it only second hand, and Runpod's pricing documentation on 21 September 2026 listed only on-demand and savings plans. Before planning around Runpod spot Pods, confirm they are still sold. Lambda's instance price list showed no spot tier on 21 September 2026.

Reserved and committed use

A reserved GPU trades flexibility for a lower rate. On demand vs reserved GPU is a question of how many hours you will use.

The hyperscalers sell one and three year terms. AWS Savings Plans commit you to "a consistent amount of usage (measured in $/hour)", with savings of "up to 66%" on Compute Savings Plans and "up to 72%" on EC2 Instance Savings Plans. Google Cloud committed use discounts offer "up to 55% discount off on-demand prices for most GPU types" and up to 65% for some. Azure Reserved VM Instances advertise "up to 72%" on one or three year terms. These are ceilings, read on 21 September 2026, and only Google's is stated for GPUs specifically.

Specialist clouds sell shorter terms. Runpod savings plans are "3 or 6 months upfront", "non-refundable", with "fixed expiration dates". Vast.ai reserved instances advertise "up to 50% discount based on commitment length", are prepaid, and have "Credits locked to the specific instance". Its API docs add that the "Maximum discount is typically 40%".

Large contracts are different again. CoreWeave's annual report for 2025, filed on 2 March 2026, describes its committed contracts as take-or-pay, "requiring payment regardless of the level of utilization", with initial terms that generally run one to six years and often involve prepayment. SemiAnalysis wrote in April 2026 that most of the rental market runs on contracts of six months or longer. Our GPU cluster pricing page covers reserved multi-node capacity.

A reserved product is not always a discount. On 21 September 2026, Lambda published per-GPU rates for its 1-Click Clusters (minimum 16 GPUs, terms from two weeks to one year) that were higher than its own on-demand 8-GPU instance rate, for both H100 and B200. Compare before assuming a reservation is cheaper.

Locking in a price can go either way. The SemiAnalysis index of one year H100 contracts rose almost 40% between October 2025 and March 2026, which rewarded early commitments. The AWS cut of June 2025 did the opposite.

Capacity blocks and calendar reservations

You book GPUs for a fixed window and pay for the whole window.

AWS Capacity Blocks for ML reserve GPU instances "on a future date to support your short duration machine learning (ML) workloads". Per AWS's documentation on 21 September 2026, a start time can be up to eight weeks ahead, a block can hold up to 64 instances, "The reservation fee is charged up front", and "Capacity Block cancellations aren't allowed". As a dated example, AWS listed a p5.48xlarge block in US East at $41.528 per hour, which it states as $5.191 per H100. AWS updates these prices "regularly based on trends in supply and demand", next in October 2026, so treat the figure as a snapshot.

One trap: AWS's docs say blocks end at 11:30 UTC, with instance termination starting at 11:00 UTC. Write your last checkpoint before then.

Google Cloud's Dynamic Workload Scheduler has two modes. Flex-start VMs "are charged based on usage, in the same way as on-demand instances". Calendar mode VMs "are tied to a reservation in which you pay for the full duration" and come only in 8 GPU shapes. Google's next price update is due in November 2026.

Marketplaces

A marketplace does not own the GPUs. Vast.ai's docs describe "a marketplace model where hosts set their own prices", billed by the second.

The risks sit in the listing. Vast.ai's own docs say higher reliability scores "typically correlate with higher prices". On-demand rentals there have a maximum duration set by the host, and "Expired instances may be deleted 48 hours after expiration." Host verification is "fully automated" and reflects machine health, not a security audit. "After spending credits, there are absolutely no refunds." Runpod's Community Cloud is a similar peer-to-peer layer, with reliability rated "Variable" in Runpod's docs.

Data transfer is the cost nobody budgets. On Vast.ai, "Data transfer costs vary by host and include both upload and download traffic." Runpod and Lambda both state they charge no ingress or egress fees. Our data egress reference compares the rest.

Checkpointing: the price of using spot

A spot discount is only real if the job survives being stopped. Each interruption costs the time since your last checkpoint, plus the restart, plus any wait for capacity to return.

  • Checkpoint on a timer, not on the warning. Thirty seconds on Azure or Google Cloud is not enough to rely on for a large checkpoint. Treat the notice as a chance to stop cleanly. AWS recommends polling for its two minute notice every 5 seconds.
  • Write to storage that outlives the instance. On Runpod, a container disk is lost when the Pod stops, a volume disk lasts until the Pod is deleted, and network volumes persist independently. Azure's default Deallocate policy keeps your disks and keeps billing them. Moving checkpoints to another provider is egress.
  • Make resume automatic. The job should find the latest checkpoint and continue unattended. Test it by killing the job yourself.
  • Set the interval from measurement. Time one checkpoint write. Pick an interval that keeps writing a small share of run time, and shorten it if interruptions are frequent. We give no number because the right interval varies with your model and storage.
  • Keep the fragile parts on-demand. Serving endpoints with users waiting do not belong on spot.

What a commitment really commits you to

Ask these before you sign or prepay anything.

  1. What exactly is committed? A spend per hour (AWS Savings Plans), a specific instance (Vast.ai locks credits to it), or named hardware for a term?
  2. Is it take-or-pay? CoreWeave's filing says its committed contracts typically are.
  3. How much is paid up front, and is any refundable? Runpod savings plans are non-refundable and AWS Capacity Blocks cannot be cancelled. Lambda's billing docs say refunds come as service credits that expire 12 months after issue.
  4. Can you cancel, shrink, or move to a newer GPU mid-term? If the contract is silent, assume not.
  5. What happens if capacity is late or hardware fails? Get the service level and credits in writing.
  6. When exactly does it start and end, and how is it invoiced? Lambda bills 1-Click Clusters in weekly increments and gives ten days to pay for a reservation.
  7. What else is billed? Storage while idle, and egress.
  8. What is your break-even utilisation? At a 40% discount, the reservation beats on-demand only if you would otherwise have run more than 60% of the committed hours.

The rule

Start on-demand and measure real utilisation for a month. Move interruption-tolerant jobs to spot once automatic resume is tested. Reserve only the baseline that was busy for more than your break-even share of hours, on the shortest term that earns most of the discount. Leave peaks on on-demand or spot. Buy a capacity block only when the start date is fixed and the code is ready to run. For a long reservation, also run the rent vs buy calculator.

Sources

Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.

Every source was accessed on 21 September 2026. A date in an entry is the publication, filing or last-update date.

Frequently asked questions

What is the difference between on-demand and reserved GPU pricing?

On-demand bills you for the time an instance runs and you can stop at any point. Reserved or committed pricing gives a lower rate in exchange for a fixed term, usually paid up front or owed whether you use the GPU or not.

How much warning do you get before a spot GPU is taken away?

AWS documents a two minute interruption notice, Azure documents 30 seconds, and Google Cloud documents a best effort shutdown period of up to 30 seconds. Vast.ai interruptible instances are stopped when outbid, which kills running processes.

Is a reserved GPU always cheaper than on-demand?

No. A reservation only saves money if the GPU is busy for enough of the committed hours, and some reserved products are priced for guaranteed capacity, not for a discount. Lambda's published cluster rate per GPU was higher than its own on-demand rate as retrieved on 21 September 2026.

Can you cancel an AWS Capacity Block?

No. AWS's documentation states that Capacity Block cancellations are not allowed, and the reservation fee is charged up front.

What workloads are safe to run on spot GPUs?

Jobs that save checkpoints on a timer to storage that outlives the instance and that restart from the latest checkpoint without a person involved. Serving endpoints with users waiting and jobs that cannot resume do not belong on spot.

Related Posts