Self-hosted LLM vs API cost calculator

Enter a month of traffic and compare paying an API per token with serving the model yourself on rented GPUs. GPU prices are live; API prices and throughput measurements come from each publisher's own page, dated below.

The short answer, at today's prices: serving Llama 3.3 70B on 2 x H100 at NVIDIA's published 2,895 output tokens per second costs $0.480 per million output tokens when the server never idles (2 GPUs at $2.50 per GPU-hour from Hyperstack, observed 2 Oct 2026, 06:24 UTC). The same model through DeepInfra costs $0.420 or Together AI costs $2.080 per million output tokens with the matching input, at prices checked on 28 September 2026. Self-hosting pays only when the servers stay busy and the cheaper API cannot do the job.

Calculator

1. Your monthly traffic
2. The API you would pay

$0.100 per million input, $0.320 per million output. Source

3. Serving it yourself

Published by NVIDIA (vendor-run): 2,895 output tokens per second across 100 concurrent requests, 1,000 input and 1,000 output tokens each, NVIDIA NIM 1.8.0, TP2. Source

Cheapest live price (Hyperstack). See all offers

Result

API, per month
$62
Self-hosted, always on, per month
$3,650
1 server x 730 hours
Self-hosted, paying only for busy hours
$48
A floor: real serving keeps headroom and idles.
API cost per million output tokens, input included
$0.620
Self-hosted cost per million output tokens, at full load
$0.480
How busy your servers would be
1.3%

At this volume the API is cheaper by $3,588 a month. One always-on server beats this API above about 5,887 million output tokens a month (at your input-to-output ratio); one server can produce about 7,609 million at the throughput used.

How the numbers work

API: input tokens times the input price plus output tokens times the output price. Batch and cache discounts are left out; if you use them, enter your effective prices as custom prices.

Self-hosted, always on: one server produces its measured output tokens per second for 730 hours a month (about 7,609 million tokens for the default server). The calculator counts the servers your monthly output needs at that rate and bills each one for every hour of the month at the GPU price times its GPU count.

Busy hours only: the hours the work takes at full load, billed at the same rate. It is a floor that assumes you could rent exactly those hours with no headroom; real services keep spare capacity for peaks.

Break-even: the monthly output at which one always-on server costs the same as the API, at your mix of input and output. If one server cannot produce that much, it never pays on cost alone at that throughput.

For the memory side of the choice, size the model with the LLM VRAM calculator; to pick serving software, see vLLM vs SGLang vs TensorRT-LLM; for traffic that comes and goes, compare serverless GPU pricing.

The throughput measurements

Every preset is a vendor-run measurement published by NVIDIA for its NIM inference microservice (table updated 1 April 2026), not a test by this site. Aggregate output tokens per second across the stated number of concurrent requests, each with the stated input and output length. Your model, engine and traffic will differ: use your own measurement when you have one.

Model and serverPrecisionEngineConcurrent requestsInput / output tokensOutput tokens/sLive $/GPU-hrLlama 3.3 70B, 2 x H100 80GBFP8NVIDIA NIM 1.8.0, TP21001,000 / 1,0002,895.49$2.50 offersLlama 3.3 70B, 2 x H200 141GBFP8NVIDIA NIM 1.8.0, TP21001,000 / 1,0003,273.32$3.43 offersLlama 3.3 70B, 8 x A100 80GBBF16NVIDIA NIM 1.8.0, TP82501,000 / 1,0002,979.05$0.67 offersLlama 3.3 70B, 2 x H100 80GBFP8NVIDIA NIM 1.8.0, TP211,000 / 1,00053.37$2.50 offers

Source: NVIDIA NIM LLM benchmarking, performance. GPU prices observed , cheapest in-stock on-demand offer at the measured card size.

API prices used

Standard on-demand prices per million tokens from each provider's own pricing page, checked on 28 September 2026. Providers change these often; check before you decide, and enter your contract price as a custom price.

ProviderModelInput $/M tokensOutput $/M tokensWeightsSourceDeepInfraLlama 3.3 70B Instruct Turbo0.10.32OpenPricing pageTogether AILlama 3.3 70B Instruct Turbo1.041.04OpenPricing pageDeepInfragpt-oss-120b0.0370.17OpenPricing pageGroqGPT OSS 120B0.150.6OpenPricing pageTogether AIGPT-OSS 120B0.150.6OpenPricing pageCerebrasgpt-oss-120b0.350.75OpenPricing pageGroqGPT OSS 20B0.0750.3OpenPricing pageDeepInfraQwen3-235B-A22B-Instruct-25070.090.55OpenPricing pageDeepInfraDeepSeek-V3.20.260.38OpenPricing pageAnthropicClaude Sonnet 5.5210ClosedPricing pageAnthropicClaude Haiku 4.515ClosedPricing pageOpenAIgpt-6-sol (prompts up to 272K tokens)210ClosedPricing pageOpenAIgpt-6-luna (short context)0.10.5ClosedPricing pageGoogleGemini 3.8 FlashPrice listed through 31 December 20260.753.75ClosedPricing pageDeepSeekdeepseek-v4-pro (peak hours)1.323.96ClosedPricing page

Prices as listed on 2026-09-28. Closed models cannot be self-hosted; they are here to price the alternative if an open model is not good enough for the job.

When self-hosting wins anyway

Cost per token is one reason among several. Serving the model yourself can still be the right call when your data cannot leave your control, when you run a fine-tuned or custom model no API offers, when you need guaranteed capacity or latency the API does not promise, or when your traffic is heavy and steady enough to keep the servers busy around the clock.

If you do self-host, rent the servers rather than buying while the volume is still uncertain: the rent vs buy calculator shows when owning starts to pay, and checking a rented GPU covers the first fifteen minutes on a new server.

Questions

Is it cheaper to self-host an LLM than to use an API?

For popular open models, often not on cost alone. At the latest price observation (2 Oct 2026, 06:24 UTC), 2 x H100 at NVIDIA's published 2,895 output tokens per second costs $0.480 per million output tokens at full load. DeepInfra charged $0.420 per million output tokens for the same model, input included, on 28 September 2026; Together AI charged $2.080 per million output tokens for the same model, input included, on 28 September 2026. Self-hosting wins when your GPUs stay busy, the API is expensive, or you need control the API cannot give.

How many tokens can one GPU server produce?

It depends on the model, precision, engine and how many requests run at once. NVIDIA's published NIM measurements for Llama 3.3 70B range from 53 output tokens per second for one request on 2 x H100 to 2,895 with 100 concurrent requests on the same server. Measure your own model before committing.

Why does the calculator use output tokens per second?

Serving benchmarks publish aggregate output tokens per second across concurrent requests, and output is what you wait for. The measurement already includes processing its own prompts, so the result holds only if your prompts are no longer relative to your answers than the benchmark's.

What does the self-hosted figure leave out?

Engineering and on-call time, storage, network transfer, monitoring and the spare capacity you keep for traffic peaks. The always-on figure assumes enough servers to carry your average load at the published throughput, running all month.