Enter a month of traffic and compare paying an API per token with serving the model yourself on rented GPUs. GPU prices are live; API prices and throughput measurements come from each publisher's own page, dated below.
The short answer, at today's prices: serving Llama 3.3 70B on 2 x H100 at NVIDIA's published 2,895 output tokens per second costs $0.480 per million output tokens when the server never idles (2 GPUs at $2.50 per GPU-hour from Hyperstack, observed 2 Oct 2026, 06:24 UTC). The same model through DeepInfra costs $0.420 or Together AI costs $2.080 per million output tokens with the matching input, at prices checked on 28 September 2026. Self-hosting pays only when the servers stay busy and the cheaper API cannot do the job.
At this volume the API is cheaper by $3,588 a month. One always-on server beats this API above about 5,887 million output tokens a month (at your input-to-output ratio); one server can produce about 7,609 million at the throughput used.
API: input tokens times the input price plus output tokens times the output price. Batch and cache discounts are left out; if you use them, enter your effective prices as custom prices.
Self-hosted, always on: one server produces its measured output tokens per second for 730 hours a month (about 7,609 million tokens for the default server). The calculator counts the servers your monthly output needs at that rate and bills each one for every hour of the month at the GPU price times its GPU count.
Busy hours only: the hours the work takes at full load, billed at the same rate. It is a floor that assumes you could rent exactly those hours with no headroom; real services keep spare capacity for peaks.
Break-even: the monthly output at which one always-on server costs the same as the API, at your mix of input and output. If one server cannot produce that much, it never pays on cost alone at that throughput.
For the memory side of the choice, size the model with the LLM VRAM calculator; to pick serving software, see vLLM vs SGLang vs TensorRT-LLM; for traffic that comes and goes, compare serverless GPU pricing.
Every preset is a vendor-run measurement published by NVIDIA for its NIM inference microservice (table updated 1 April 2026), not a test by this site. Aggregate output tokens per second across the stated number of concurrent requests, each with the stated input and output length. Your model, engine and traffic will differ: use your own measurement when you have one.
Source: NVIDIA NIM LLM benchmarking, performance. GPU prices observed , cheapest in-stock on-demand offer at the measured card size.
Standard on-demand prices per million tokens from each provider's own pricing page, checked on 28 September 2026. Providers change these often; check before you decide, and enter your contract price as a custom price.
Prices as listed on 2026-09-28. Closed models cannot be self-hosted; they are here to price the alternative if an open model is not good enough for the job.
Cost per token is one reason among several. Serving the model yourself can still be the right call when your data cannot leave your control, when you run a fine-tuned or custom model no API offers, when you need guaranteed capacity or latency the API does not promise, or when your traffic is heavy and steady enough to keep the servers busy around the clock.
If you do self-host, rent the servers rather than buying while the volume is still uncertain: the rent vs buy calculator shows when owning starts to pay, and checking a rented GPU covers the first fifteen minutes on a new server.
For popular open models, often not on cost alone. At the latest price observation (2 Oct 2026, 06:24 UTC), 2 x H100 at NVIDIA's published 2,895 output tokens per second costs $0.480 per million output tokens at full load. DeepInfra charged $0.420 per million output tokens for the same model, input included, on 28 September 2026; Together AI charged $2.080 per million output tokens for the same model, input included, on 28 September 2026. Self-hosting wins when your GPUs stay busy, the API is expensive, or you need control the API cannot give.
It depends on the model, precision, engine and how many requests run at once. NVIDIA's published NIM measurements for Llama 3.3 70B range from 53 output tokens per second for one request on 2 x H100 to 2,895 with 100 concurrent requests on the same server. Measure your own model before committing.
Serving benchmarks publish aggregate output tokens per second across concurrent requests, and output is what you wait for. The measurement already includes processing its own prompts, so the result holds only if your prompts are no longer relative to your answers than the benchmark's.
Engineering and on-call time, storage, network transfer, monitoring and the spare capacity you keep for traffic peaks. The always-on figure assumes enough servers to carry your average load at the published throughput, running all month.