Start with vLLM or SGLang for shared NVIDIA or AMD serving, and TensorRT-LLM for NVIDIA's own optimized stack, allowing for a build step in older workflows but following NVIDIA's migration guide, updated 21 September 2026, which says the current backend removes it. Choose Ollama or llama.cpp for local, CPU or Apple setups. These picks follow each project's documented platform support; they are not speed rankings. Avoid TGI for a new deployment: its maintenance-mode change was merged on 11 December 2025 and Hugging Face archived the repository on 21 March 2026.
Choose the model checkpoint first, then its precision, then the serving software and GPU. Treat the recommendations here as starting points. The projects' September documentation establishes features, not a universal winner. If you are still separating training requirements from serving requirements, start with AI training vs inference.
The feature matrix, dated 28 September 2026
The table attributes capabilities to each project's documentation, retrieved 28 September 2026, and versions to its release pages. Versions shown are the latest stable tags as of 28 September 2026. Quantization entries are examples, not promises that every listed format works on every backend. An unknown multi-node entry means this comparison does not establish support.
| Engine | Maintainer and licence | Stable version and release date | Key techniques | Hardware documented | Quantization examples | Multi-GPU; multi-node | OpenAI-compatible API |
|---|---|---|---|---|---|---|---|
| vLLM | vLLM community, originated at Berkeley; Apache-2.0 | 0.30.0; 22 Sep 2026 | PagedAttention, continuous batching, chunked prefill, prefix caching, speculation | NVIDIA; AMD ROCm; CPUs; Apple via vLLM-Metal | FP8, MXFP4, NVFP4, GPTQ, AWQ | Tensor/pipeline/data/expert/context; documented multi-node | Yes |
| SGLang | Community under LMSYS; Apache-2.0 | 0.5.20; 18 Sep 2026 | RadixAttention, paged attention, continuous batching, speculation, HiCache | NVIDIA; AMD MI300/MI355; Intel Xeon | FP8, AWQ, gptq_marlin, modelopt_fp4, Petit NVFP4 | Tensor/pipeline/expert/data; distributed clusters | Yes |
| TensorRT-LLM | NVIDIA; Apache-2.0 with separately licensed portions | 1.2.1; 20 Apr 2026 | PyTorch-native execution, in-flight batching, paged attention, speculation, disaggregation | NVIDIA GPUs; check model and release | NVFP4, MXFP4, FP8, AWQ/GPTQ by model | Single/multi-GPU and multi-node LLM API | OpenAI-style chat endpoint |
| llama.cpp | GGML team and community, technical leadership retained after joining Hugging Face; MIT | 0.5.0; 23 Sep 2026 | CPU/GPU offload, continuous batching, speculation | Apple Metal; x86 CPUs; NVIDIA CUDA; AMD HIP; Vulkan; SYCL | GGUF integer quantization; NVFP4/MXFP4 CUDA paths | Layer split; experimental tensor split; remote RPC proof of concept | Chat, Responses, embeddings |
| Ollama | Ollama; MIT | 0.34.4; 23 Sep 2026 | Model management, parallel requests, backend-dependent Flash Attention | NVIDIA; listed AMD GPUs; Apple Metal; CPU; Vulkan | Prepared GGUF; f16/q8_0/q4_0 KV cache; MLX NVFP4 | Host multi-GPU placement; multi-node unknown | Explicit API subset |
| TGI | Hugging Face, archived; Apache-2.0 | 3.3.7; 19 Dec 2025 | Continuous batching, Flash Attention, PagedAttention, speculation | Named NVIDIA GPUs; tested AMD MI210/MI250/MI300; llama.cpp backend for CPU/GPU | GPTQ, AWQ, bitsandbytes, EETQ, Marlin, EXL2, FP8 | Tensor parallelism; multi-node unknown | Chat Completions Messages API |
Do not equate the stable tag with every feature in rolling documentation. NVIDIA also lists TensorRT-LLM 1.3.0rc28, released 23 September 2026, as a prerelease. Pin the container and checkpoint together, and check that release's model support before reserving hardware.
Match the engine to the work
vLLM: the first shared endpoint to try
Start here when you want a baseline for concurrent serving. The vLLM project's documentation, retrieved 28 September 2026, lists continuous batching, prefix caching and several parallelism modes. Its scaling guide recommends one GPU when the model fits, tensor parallelism within a node when it does not, and combined tensor/pipeline parallelism beyond one node.
Check the checkpoint before assuming compatibility. vLLM's GGUF support now sits in a separate plugin, and its documentation calls the GGUF path experimental and under-optimized, with possible incompatibilities with other features. vLLM's 0.30.0 release also removes GPTQ activation ordering through g_idx. For repeated-document applications, vLLM's 20 September 2026 prefix-cache guide says reuse saves prefill computation but does not shorten new-token decoding. That distinction should shape your latency target.
SGLang vs vLLM: compare prefix reuse on your traffic
SGLang deserves a comparison when conversations share long prefixes. SGLang's repository page, retrieved 28 September 2026, lists RadixAttention alongside paged attention, so those are not opposing design choices. SGLang's HiCache documentation, retrieved 28 September 2026, describes GPU, host-memory and optional storage tiers for reuse.
The same documentation says host memory remains private to each inference instance. Cross-instance reuse works through a suitably configured shared storage backend. Do not plan capacity as though several servers automatically share one RAM pool. SGLang's quantization guide also replaces plain GPU GPTQ with gptq_marlin and recommends offline quantization. Confirm the exact loading method before moving an existing checkpoint.
TensorRT-LLM: choose the NVIDIA path deliberately
TensorRT-LLM is a candidate for supported NVIDIA deployments and NVIDIA-specific kernels, going by NVIDIA's overview; that is not a measured lead over vLLM. NVIDIA's migration guide, updated 21 September 2026, says the TensorRT engine backend has been removed. The current path loads Hugging Face checkpoints directly and eliminates separate conversion and trtllm-build.
That ends the separate conversion and build steps that older instructions describe. Keep older instructions tied to their release. NVIDIA's quantization documentation, retrieved 28 September 2026, remains model-specific: its NVFP4 KV-cache path requires offline ModelOpt quantization and FP8 weights/activations. A format name alone is not a deployment recipe.
vLLM vs Ollama: decide how much serving control you need
For an individual developer, start with Ollama's model-management workflow. Ollama's documentation, retrieved 28 September 2026, describes a CLI and REST API, GGUF import, concurrent requests and placement across GPUs when one GPU cannot hold the model. Calling it strictly single-user would be wrong.
For a shared rented endpoint, my Ollama vs vLLM recommendation is to start the comparison with vLLM. Ollama's FAQ says request parallelism multiplied by context length increases memory needs, and its API documentation promises only a subset of OpenAI compatibility. Exercise the exact routes your application uses. Ollama's import guide also says it does not quantize GGUF files during import, so prepare them beforehand.
llama.cpp vs vLLM: portability or a shared serving baseline
Choose llama.cpp when CPU/GPU offload or Apple deployment is central. That follows the GGML project's documented hardware goals. Its server documentation, retrieved 28 September 2026, includes continuous batching and parallel decoding, so local use is a recommendation, not a concurrency limit.
GGML's multi-GPU documentation calls layer splitting the most compatible choice and tensor splitting experimental. Its RPC documentation calls remote-device execution a fragile, insecure proof of concept and says never to expose it on an open network. Do not turn that feature into a production cluster plan without addressing those explicit limits.
TGI: plan the migration around the endpoint
Keep an existing TGI deployment only with an explicit maintenance plan. Hugging Face's endpoint guidance, retrieved 28 September 2026, recommends vLLM or SGLang. When a Hugging Face TGI endpoint moves to vLLM, Hugging Face's guide requires creating a new endpoint before switching traffic. Preserve your request and response checks during that move; an engine change should not silently change application behavior.
Triton Inference Server and Dynamo: serving layers
NVIDIA's documentation, retrieved 28 September 2026, describes Triton Inference Server as a general model server with dynamic batching and concurrent model execution. Its TensorRT-LLM backend documents a PyTorch LLM-API path without engine compilation. The separate OpenAI frontend supports TensorRT-LLM orchestrator mode, but not leader mode, for model parallelism.
NVIDIA describes Dynamo as orchestration above vLLM, SGLang and TensorRT-LLM, with cache-aware routing and independently scalable prefill/decode pools. Its README says an engine alone is probably enough for one model on one GPU. Start there; add orchestration when you have a concrete routing or scaling requirement.
Published results are workload comparisons
A benchmark result belongs to the model, precision, request mix and software versions it was run with. Use a published vLLM benchmark to form a shortlist, then compare your own latency and throughput requirements under the same request distribution. None of the following measurements were made by this site.
| Publication and label | Model, precision, GPU and setup | Reported finding and boundary |
|---|---|---|
| LMSYS, 25 July 2024; PROJECT BENCHMARK / VENDOR CLAIM | Llama-8B, BF16, one A100. SGLang v0.2 study; vLLM 0.5.2 defaults; TensorRT-LLM 0.10.0 recommended arguments and tuned batches. Offline synthetic and ShareGPT workloads, 1K to 6K requests together; prefix caching and speculative decoding disabled. OpenAI interfaces for SGLang/vLLM, Triton for TensorRT-LLM; equal output lengths enforced. | SGLang and TensorRT-LLM reached "up to 5000 tokens per second" in short-input tests, ahead of vLLM. Output throughput was measured over total duration. Synthetic Input-512-Output-1024 sampled independent uniform lengths from 1 to each limit. LMSYS corrected a short-input generation bias on 26 July. Historical, not a current ordering. |
| SemiAnalysis, 9 October 2025 article describing 7 October results; INDEPENDENT BENCHMARK | Llama 3.3 70B FP8, MI300X versus H100 using vLLM; MI300X TP1 with ROCm 7.0. Reasoning workload: 1024 input/8192 output tokens; input lengths randomized to 80% to 100%; random tokens avoid prefix reuse. Infinite offered rate with bounded concurrency; concurrency/parallelism sweeps. | SemiAnalysis describes strong MI300X performance against H100, particularly at "20 to 30 tok/s/user." This compares hardware configurations, not vLLM with SGLang; InferenceMAX v1 selected only one of vLLM or SGLang as the default engine for each model. |
| NVIDIA, 9 September 2025; VENDOR CLAIM / MLPerf SUBMISSION | DeepSeek-R1, TensorRT-LLM, 72-GPU GB300 NVL72 versus 72-GPU GB200 NVL72. Most weights converted from FP8 to NVFP4 with Model Optimizer; FP8 KV cache. MLPerf Inference v5.1 Closed offline, entries 5.1-0072/0071. Expert parallelism for MoE, attention data parallelism, Attention Data Parallelism Balance and decode-only CUDA Graphs. | NVIDIA reports 5,842 tokens/sec/GPU on GB300 NVL72 versus 4,024 on GB200 NVL72. These are NVIDIA's normalized figures, not MLPerf's primary metric. This is a system comparison, not an isolated engine comparison. |
For your comparison, hold the checkpoint, precision, input/output lengths and concurrency constant. Baseten's performance guide, retrieved 28 September 2026, recommends production-like contents because they affect prefix-cache hits and speculative acceptance. It also recommends matching temperature and reasoning effort. Record time to first token separately from subsequent token latency, and retain the exact image and command with the result.
Before renting for a longer run, write down your acceptance criteria: the slowest first response you will tolerate, the generation pace each user needs, and the number of simultaneous conversations. Include a cold start and a repeated conversation in your trial. Keep those results separate. A warm cache should not hide an unacceptable first request. Reject configurations that miss your application target even if their aggregate token count looks better.
Fit the checkpoint to the rented GPU
Use the LLM VRAM calculator before selecting a server. Compare the candidate families below, then check the engine's model-specific precision path. Leave room in your plan for concurrent contexts rather than budgeting only for weights.
| Spec | H100 | H200 | B200 | L40S | RTX 4090 | MI300X |
|---|---|---|---|---|---|---|
| VRAM | 80 to 94 GB | 141 GB | 180 to 192 GB | 48 GB | 24 GB | 192 GB |
| Memory type | HBM3 | HBM3e | HBM3e | GDDR6 | GDDR6X | HBM3 |
| Memory bandwidth | 3,350 GB/s | 4,800 GB/s | 8,000 GB/s | 864 GB/s | 1,008 GB/s | 5,300 GB/s |
| FP8 (dense) | 1,979 TFLOPS | 1,979 TFLOPS | 4,500 TFLOPS | 733 TFLOPS | 330.3 TFLOPS | 2,614.9 TFLOPS |
| FP8 (with sparsity) | 3,958 TFLOPS | 3,958 TFLOPS | 9,000 TFLOPS | 1,466 TFLOPS | 660.6 TFLOPS | 5,229.8 TFLOPS |
| FP4 (dense) | Not published | Not published | 9,000 TFLOPS | Not published | Not published | Not published |
| FP4 (with sparsity) | Not published | Not published | 18,000 TFLOPS | Not published | Not published | Not published |
FP8 and FP4 support must match across checkpoint, engine kernel and GPU if you want native acceleration; accepting a checkpoint does not establish native arithmetic. SGLang's quantization documentation, retrieved 28 September 2026, explicitly distinguishes its Marlin W4A16 fallback from native FP4 backends. Read FP8 vs FP16 vs BF16 before choosing a format.
For MI300X, check the ROCm deployment path against ROCm vs CUDA. For H100, H200, B200, L40S or RTX 4090, still check the selected kernel rather than treating NVIDIA support as universal. If the model needs several GPUs, choose the split using tensor, pipeline and data parallelism before comparing server listings.
Where to run the engine
Separate a software template from a managed model endpoint. The provider documents below describe deployment paths; they do not establish current GPU inventory. Pin the image even when the deployment starts with a button. For intermittent traffic, compare the operating model in serverless GPU pricing.
| Provider documentation and date | Documented deployment path | What to check |
|---|---|---|
| RunPod, 13 and 25 September 2026; TensorRT-LLM guide retrieved 28 September 2026 | Prebuilt vLLM Serverless worker; SGLang Quick Deploy; TensorRT-LLM through custom container and handler | RunPod recommends pinning the worker image; its vLLM worker exposes tensor parallelism |
| Vast.ai, 26 February 2025 | vLLM template using vllm/vllm-openai | Its stable-endpoint example requires a direct forwarded port and static IP |
| Nebius, retrieved 28 September 2026 | Marketplace vLLM on a VM or Managed Kubernetes | The deployment is self-managed |
| DigitalOcean, 14 June and 20 July 2026 | vLLM Single GPU image for H100/H200; Ollama with Open WebUI application | Catalog inventories list vLLM 0.19.1 and Ollama 0.10.2; compare with upstream before relying on newer features |
| AWS, 18 September 2026 | Managed SageMaker endpoints with vLLM and SGLang Deep Learning Containers | Match the container to the intended engine |
| Lambda, retrieved 28 September 2026 | Self-managed vLLM on a 1-Click Cluster using Ray and InfiniBand | This is a cluster tutorial, not a managed model endpoint |
The live table below shows rental comparisons. Use it after selecting the engine and memory requirement, then confirm the deployment image matches your intended version.
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| H100 | $2.59 | QuantaCloud | 7 |
| H200 | $3.43 | QuantaCloud | 7 |
| L40S | $0.97 | Massed Compute | 5 |
| RTX 4090 | $0.43 | Vast.ai | 3 |
| MI300X | none in stock | ||
For a template, review who owns upgrades, restarts and endpoint configuration. For a managed endpoint, ask which engine settings you can change and how you can pin or roll back a version. Keep a known working configuration before changing the image. Choose the deployment whose operating responsibilities your team can actually take on.
The starting decision
These are editorial starting recommendations based on the documented limits above, not measured performance rankings.
| Situation | Engine to start with |
|---|---|
| Shared NVIDIA or AMD endpoint, no specialized requirement yet | vLLM |
| Long conversations or repeated prefixes dominate | Compare SGLang with the same vLLM workload |
| Supported NVIDIA model and a need for NVIDIA-specific kernels | TensorRT-LLM, pinned to a specific release |
| Individual developer wants model management | Ollama |
| CPU/GPU offload, Apple deployment or direct GGUF control | llama.cpp |
| Existing TGI endpoint needs a future maintenance path | Create and validate a vLLM or SGLang replacement |
| Multiple model frameworks or cluster routing requirements | Evaluate Triton or Dynamo around the selected engine |
Start with vLLM for the rented shared endpoint. Switch only when a supported feature or a workload-matched comparison gives you a concrete reason.
Sources
- GitHub: vllm-project/vllm, retrieved 2026-09-28.
- GitHub: vllm-project/vllm releases, 2026-09-22; 2026-09-09.
- vLLM: GPU installation, retrieved 2026-09-28.
- vLLM: GGUF quantization, retrieved 2026-09-28.
- vLLM: parallelism and scaling, retrieved 2026-09-28.
- vLLM: automatic prefix caching, 2026-09-20.
- GitHub: sgl-project/sglang, retrieved 2026-09-28.
- GitHub: sgl-project/sglang releases, 2026-09-18.
- SGLang: quantization, retrieved 2026-09-28.
- SGLang: HiCache design, retrieved 2026-09-28.
- GitHub: NVIDIA/TensorRT-LLM, retrieved 2026-09-28.
- GitHub: NVIDIA/TensorRT-LLM licence file, retrieved 2026-09-28.
- GitHub: NVIDIA/TensorRT-LLM v1.2.1 release, 2026-04-20.
- GitHub: NVIDIA/TensorRT-LLM releases, 2026-09-23; retrieved 2026-09-28.
- NVIDIA TensorRT-LLM: overview, retrieved 2026-09-28.
- NVIDIA TensorRT-LLM: quantization, retrieved 2026-09-28.
- NVIDIA TensorRT-LLM: quick start guide, retrieved 2026-09-28.
- NVIDIA TensorRT-LLM: TensorRT backend removal, 2026-09-21.
- GitHub: ggml-org/llama.cpp, retrieved 2026-09-28.
- GitHub: ggml-org/llama.cpp v0.5.0 release, 2026-09-23.
- Hugging Face: GGML joins Hugging Face, 2026-02-20.
- Hugging Face Inference Endpoints: llama.cpp engine, retrieved 2026-09-28.
- GitHub: llama.cpp server documentation, retrieved 2026-09-28.
- GitHub: llama.cpp multi-GPU documentation, retrieved 2026-09-28.
- GitHub: llama.cpp RPC documentation, retrieved 2026-09-28.
- GitHub: llama.cpp build documentation, retrieved 2026-09-28.
- GitHub: ollama/ollama, retrieved 2026-09-28.
- GitHub: ollama/ollama releases, 2026-09-23; 2026-09-28.
- GitHub: Ollama README, retrieved 2026-09-28.
- GitHub: Ollama licence file, retrieved 2026-09-28.
- Ollama: GPU support, retrieved 2026-09-28.
- Ollama: importing models, retrieved 2026-09-28.
- Ollama: OpenAI compatibility, retrieved 2026-09-28.
- Ollama: FAQ, retrieved 2026-09-28.
- Ollama: MLX performance, 2026-06-11.
- GitHub: huggingface/text-generation-inference, retrieved 2026-09-28.
- GitHub: huggingface/text-generation-inference releases, 2025-12-19.
- GitHub: huggingface/text-generation-inference pull request 3344, 2025-12-11.
- Hugging Face TGI: documentation index, retrieved 2026-09-28.
- Hugging Face TGI: speculation, retrieved 2026-09-28.
- Hugging Face TGI: quantization, retrieved 2026-09-28.
- Hugging Face TGI: NVIDIA installation, retrieved 2026-09-28.
- Hugging Face TGI: AMD installation, retrieved 2026-09-28.
- Hugging Face TGI: llama.cpp backend, retrieved 2026-09-28.
- Hugging Face TGI: Messages API, retrieved 2026-09-28.
- Hugging Face Inference Endpoints: TGI engine, retrieved 2026-09-28.
- GitHub: triton-inference-server/server, retrieved 2026-09-28.
- NVIDIA Triton: TensorRT-LLM backend README, retrieved 2026-09-28.
- NVIDIA Triton: OpenAI client guide, retrieved 2026-09-28.
- GitHub: ai-dynamo/dynamo, retrieved 2026-09-28.
- LMSYS: SGLang Llama 3 serving comparison, 2024-07-25.
- SemiAnalysis: InferenceMAX open-source inference benchmarking, 2025-10-09.
- NVIDIA: Blackwell Ultra sets new inference records in MLPerf debut, 2025-09-09.
- Baseten: performance benchmarking and load testing, retrieved 2026-09-28.
- RunPod: run vLLM on RunPod Serverless, 2026-09-13.
- RunPod: Serverless LLM guide, 2026-09-25.
- RunPod: vLLM vs TensorRT-LLM, retrieved 2026-09-28.
- Vast.ai: serving DeepSeek models with vLLM and LangChain, 2025-02-26.
- Nebius Marketplace: vLLM, retrieved 2026-09-28.
- DigitalOcean Marketplace: 1-Click Inference Ready Single GPU, 2026-06-14.
- DigitalOcean Marketplace: Ollama with Open WebUI, 2026-07-20.
- AWS: Amazon SageMaker inference 2026 year-to-date launches in review, 2026-09-18.
- Lambda: serving Llama 3.1 405B, retrieved 2026-09-28.