vLLM vs SGLang vs TensorRT-LLM: Choose by Workload

Choose an inference engine for your rented GPU using dated feature support, published benchmark setups, deployment options and practical workload rules.

By Faiz Ahmed•
•15 min read

Start with vLLM or SGLang for shared NVIDIA or AMD serving, and TensorRT-LLM for NVIDIA's own optimized stack, allowing for a build step in older workflows but following NVIDIA's migration guide, updated 21 September 2026, which says the current backend removes it. Choose Ollama or llama.cpp for local, CPU or Apple setups. These picks follow each project's documented platform support; they are not speed rankings. Avoid TGI for a new deployment: its maintenance-mode change was merged on 11 December 2025 and Hugging Face archived the repository on 21 March 2026.

Choose the model checkpoint first, then its precision, then the serving software and GPU. Treat the recommendations here as starting points. The projects' September documentation establishes features, not a universal winner. If you are still separating training requirements from serving requirements, start with AI training vs inference.

The feature matrix, dated 28 September 2026

The table attributes capabilities to each project's documentation, retrieved 28 September 2026, and versions to its release pages. Versions shown are the latest stable tags as of 28 September 2026. Quantization entries are examples, not promises that every listed format works on every backend. An unknown multi-node entry means this comparison does not establish support.

EngineMaintainer and licenceStable version and release dateKey techniquesHardware documentedQuantization examplesMulti-GPU; multi-nodeOpenAI-compatible API
vLLMvLLM community, originated at Berkeley; Apache-2.00.30.0; 22 Sep 2026PagedAttention, continuous batching, chunked prefill, prefix caching, speculationNVIDIA; AMD ROCm; CPUs; Apple via vLLM-MetalFP8, MXFP4, NVFP4, GPTQ, AWQTensor/pipeline/data/expert/context; documented multi-nodeYes
SGLangCommunity under LMSYS; Apache-2.00.5.20; 18 Sep 2026RadixAttention, paged attention, continuous batching, speculation, HiCacheNVIDIA; AMD MI300/MI355; Intel XeonFP8, AWQ, gptq_marlin, modelopt_fp4, Petit NVFP4Tensor/pipeline/expert/data; distributed clustersYes
TensorRT-LLMNVIDIA; Apache-2.0 with separately licensed portions1.2.1; 20 Apr 2026PyTorch-native execution, in-flight batching, paged attention, speculation, disaggregationNVIDIA GPUs; check model and releaseNVFP4, MXFP4, FP8, AWQ/GPTQ by modelSingle/multi-GPU and multi-node LLM APIOpenAI-style chat endpoint
llama.cppGGML team and community, technical leadership retained after joining Hugging Face; MIT0.5.0; 23 Sep 2026CPU/GPU offload, continuous batching, speculationApple Metal; x86 CPUs; NVIDIA CUDA; AMD HIP; Vulkan; SYCLGGUF integer quantization; NVFP4/MXFP4 CUDA pathsLayer split; experimental tensor split; remote RPC proof of conceptChat, Responses, embeddings
OllamaOllama; MIT0.34.4; 23 Sep 2026Model management, parallel requests, backend-dependent Flash AttentionNVIDIA; listed AMD GPUs; Apple Metal; CPU; VulkanPrepared GGUF; f16/q8_0/q4_0 KV cache; MLX NVFP4Host multi-GPU placement; multi-node unknownExplicit API subset
TGIHugging Face, archived; Apache-2.03.3.7; 19 Dec 2025Continuous batching, Flash Attention, PagedAttention, speculationNamed NVIDIA GPUs; tested AMD MI210/MI250/MI300; llama.cpp backend for CPU/GPUGPTQ, AWQ, bitsandbytes, EETQ, Marlin, EXL2, FP8Tensor parallelism; multi-node unknownChat Completions Messages API

Do not equate the stable tag with every feature in rolling documentation. NVIDIA also lists TensorRT-LLM 1.3.0rc28, released 23 September 2026, as a prerelease. Pin the container and checkpoint together, and check that release's model support before reserving hardware.

Match the engine to the work

vLLM: the first shared endpoint to try

Start here when you want a baseline for concurrent serving. The vLLM project's documentation, retrieved 28 September 2026, lists continuous batching, prefix caching and several parallelism modes. Its scaling guide recommends one GPU when the model fits, tensor parallelism within a node when it does not, and combined tensor/pipeline parallelism beyond one node.

Check the checkpoint before assuming compatibility. vLLM's GGUF support now sits in a separate plugin, and its documentation calls the GGUF path experimental and under-optimized, with possible incompatibilities with other features. vLLM's 0.30.0 release also removes GPTQ activation ordering through g_idx. For repeated-document applications, vLLM's 20 September 2026 prefix-cache guide says reuse saves prefill computation but does not shorten new-token decoding. That distinction should shape your latency target.

SGLang vs vLLM: compare prefix reuse on your traffic

SGLang deserves a comparison when conversations share long prefixes. SGLang's repository page, retrieved 28 September 2026, lists RadixAttention alongside paged attention, so those are not opposing design choices. SGLang's HiCache documentation, retrieved 28 September 2026, describes GPU, host-memory and optional storage tiers for reuse.

The same documentation says host memory remains private to each inference instance. Cross-instance reuse works through a suitably configured shared storage backend. Do not plan capacity as though several servers automatically share one RAM pool. SGLang's quantization guide also replaces plain GPU GPTQ with gptq_marlin and recommends offline quantization. Confirm the exact loading method before moving an existing checkpoint.

TensorRT-LLM: choose the NVIDIA path deliberately

TensorRT-LLM is a candidate for supported NVIDIA deployments and NVIDIA-specific kernels, going by NVIDIA's overview; that is not a measured lead over vLLM. NVIDIA's migration guide, updated 21 September 2026, says the TensorRT engine backend has been removed. The current path loads Hugging Face checkpoints directly and eliminates separate conversion and trtllm-build.

That ends the separate conversion and build steps that older instructions describe. Keep older instructions tied to their release. NVIDIA's quantization documentation, retrieved 28 September 2026, remains model-specific: its NVFP4 KV-cache path requires offline ModelOpt quantization and FP8 weights/activations. A format name alone is not a deployment recipe.

vLLM vs Ollama: decide how much serving control you need

For an individual developer, start with Ollama's model-management workflow. Ollama's documentation, retrieved 28 September 2026, describes a CLI and REST API, GGUF import, concurrent requests and placement across GPUs when one GPU cannot hold the model. Calling it strictly single-user would be wrong.

For a shared rented endpoint, my Ollama vs vLLM recommendation is to start the comparison with vLLM. Ollama's FAQ says request parallelism multiplied by context length increases memory needs, and its API documentation promises only a subset of OpenAI compatibility. Exercise the exact routes your application uses. Ollama's import guide also says it does not quantize GGUF files during import, so prepare them beforehand.

llama.cpp vs vLLM: portability or a shared serving baseline

Choose llama.cpp when CPU/GPU offload or Apple deployment is central. That follows the GGML project's documented hardware goals. Its server documentation, retrieved 28 September 2026, includes continuous batching and parallel decoding, so local use is a recommendation, not a concurrency limit.

GGML's multi-GPU documentation calls layer splitting the most compatible choice and tensor splitting experimental. Its RPC documentation calls remote-device execution a fragile, insecure proof of concept and says never to expose it on an open network. Do not turn that feature into a production cluster plan without addressing those explicit limits.

TGI: plan the migration around the endpoint

Keep an existing TGI deployment only with an explicit maintenance plan. Hugging Face's endpoint guidance, retrieved 28 September 2026, recommends vLLM or SGLang. When a Hugging Face TGI endpoint moves to vLLM, Hugging Face's guide requires creating a new endpoint before switching traffic. Preserve your request and response checks during that move; an engine change should not silently change application behavior.

Triton Inference Server and Dynamo: serving layers

NVIDIA's documentation, retrieved 28 September 2026, describes Triton Inference Server as a general model server with dynamic batching and concurrent model execution. Its TensorRT-LLM backend documents a PyTorch LLM-API path without engine compilation. The separate OpenAI frontend supports TensorRT-LLM orchestrator mode, but not leader mode, for model parallelism.

NVIDIA describes Dynamo as orchestration above vLLM, SGLang and TensorRT-LLM, with cache-aware routing and independently scalable prefill/decode pools. Its README says an engine alone is probably enough for one model on one GPU. Start there; add orchestration when you have a concrete routing or scaling requirement.

Published results are workload comparisons

A benchmark result belongs to the model, precision, request mix and software versions it was run with. Use a published vLLM benchmark to form a shortlist, then compare your own latency and throughput requirements under the same request distribution. None of the following measurements were made by this site.

Publication and labelModel, precision, GPU and setupReported finding and boundary
LMSYS, 25 July 2024; PROJECT BENCHMARK / VENDOR CLAIMLlama-8B, BF16, one A100. SGLang v0.2 study; vLLM 0.5.2 defaults; TensorRT-LLM 0.10.0 recommended arguments and tuned batches. Offline synthetic and ShareGPT workloads, 1K to 6K requests together; prefix caching and speculative decoding disabled. OpenAI interfaces for SGLang/vLLM, Triton for TensorRT-LLM; equal output lengths enforced.SGLang and TensorRT-LLM reached "up to 5000 tokens per second" in short-input tests, ahead of vLLM. Output throughput was measured over total duration. Synthetic Input-512-Output-1024 sampled independent uniform lengths from 1 to each limit. LMSYS corrected a short-input generation bias on 26 July. Historical, not a current ordering.
SemiAnalysis, 9 October 2025 article describing 7 October results; INDEPENDENT BENCHMARKLlama 3.3 70B FP8, MI300X versus H100 using vLLM; MI300X TP1 with ROCm 7.0. Reasoning workload: 1024 input/8192 output tokens; input lengths randomized to 80% to 100%; random tokens avoid prefix reuse. Infinite offered rate with bounded concurrency; concurrency/parallelism sweeps.SemiAnalysis describes strong MI300X performance against H100, particularly at "20 to 30 tok/s/user." This compares hardware configurations, not vLLM with SGLang; InferenceMAX v1 selected only one of vLLM or SGLang as the default engine for each model.
NVIDIA, 9 September 2025; VENDOR CLAIM / MLPerf SUBMISSIONDeepSeek-R1, TensorRT-LLM, 72-GPU GB300 NVL72 versus 72-GPU GB200 NVL72. Most weights converted from FP8 to NVFP4 with Model Optimizer; FP8 KV cache. MLPerf Inference v5.1 Closed offline, entries 5.1-0072/0071. Expert parallelism for MoE, attention data parallelism, Attention Data Parallelism Balance and decode-only CUDA Graphs.NVIDIA reports 5,842 tokens/sec/GPU on GB300 NVL72 versus 4,024 on GB200 NVL72. These are NVIDIA's normalized figures, not MLPerf's primary metric. This is a system comparison, not an isolated engine comparison.

For your comparison, hold the checkpoint, precision, input/output lengths and concurrency constant. Baseten's performance guide, retrieved 28 September 2026, recommends production-like contents because they affect prefix-cache hits and speculative acceptance. It also recommends matching temperature and reasoning effort. Record time to first token separately from subsequent token latency, and retain the exact image and command with the result.

Before renting for a longer run, write down your acceptance criteria: the slowest first response you will tolerate, the generation pace each user needs, and the number of simultaneous conversations. Include a cold start and a repeated conversation in your trial. Keep those results separate. A warm cache should not hide an unacceptable first request. Reject configurations that miss your application target even if their aggregate token count looks better.

Fit the checkpoint to the rented GPU

Use the LLM VRAM calculator before selecting a server. Compare the candidate families below, then check the engine's model-specific precision path. Leave room in your plan for concurrent contexts rather than budgeting only for weights.

SpecH100H200B200L40SRTX 4090MI300X
VRAM80 to 94 GB141 GB180 to 192 GB48 GB24 GB192 GB
Memory typeHBM3HBM3eHBM3eGDDR6GDDR6XHBM3
Memory bandwidth3,350 GB/s4,800 GB/s8,000 GB/s864 GB/s1,008 GB/s5,300 GB/s
FP8 (dense)1,979 TFLOPS1,979 TFLOPS4,500 TFLOPS733 TFLOPS330.3 TFLOPS2,614.9 TFLOPS
FP8 (with sparsity)3,958 TFLOPS3,958 TFLOPS9,000 TFLOPS1,466 TFLOPS660.6 TFLOPS5,229.8 TFLOPS
FP4 (dense)Not publishedNot published9,000 TFLOPSNot publishedNot publishedNot published
FP4 (with sparsity)Not publishedNot published18,000 TFLOPSNot publishedNot publishedNot published
Figures from the vendor datasheets: H100, H200, B200, L40S, RTX 4090, MI300X, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

FP8 and FP4 support must match across checkpoint, engine kernel and GPU if you want native acceleration; accepting a checkpoint does not establish native arithmetic. SGLang's quantization documentation, retrieved 28 September 2026, explicitly distinguishes its Marlin W4A16 fallback from native FP4 backends. Read FP8 vs FP16 vs BF16 before choosing a format.

For MI300X, check the ROCm deployment path against ROCm vs CUDA. For H100, H200, B200, L40S or RTX 4090, still check the selected kernel rather than treating NVIDIA support as universal. If the model needs several GPUs, choose the split using tensor, pipeline and data parallelism before comparing server listings.

Where to run the engine

Separate a software template from a managed model endpoint. The provider documents below describe deployment paths; they do not establish current GPU inventory. Pin the image even when the deployment starts with a button. For intermittent traffic, compare the operating model in serverless GPU pricing.

Provider documentation and dateDocumented deployment pathWhat to check
RunPod, 13 and 25 September 2026; TensorRT-LLM guide retrieved 28 September 2026Prebuilt vLLM Serverless worker; SGLang Quick Deploy; TensorRT-LLM through custom container and handlerRunPod recommends pinning the worker image; its vLLM worker exposes tensor parallelism
Vast.ai, 26 February 2025vLLM template using vllm/vllm-openaiIts stable-endpoint example requires a direct forwarded port and static IP
Nebius, retrieved 28 September 2026Marketplace vLLM on a VM or Managed KubernetesThe deployment is self-managed
DigitalOcean, 14 June and 20 July 2026vLLM Single GPU image for H100/H200; Ollama with Open WebUI applicationCatalog inventories list vLLM 0.19.1 and Ollama 0.10.2; compare with upstream before relying on newer features
AWS, 18 September 2026Managed SageMaker endpoints with vLLM and SGLang Deep Learning ContainersMatch the container to the intended engine
Lambda, retrieved 28 September 2026Self-managed vLLM on a 1-Click Cluster using Ray and InfiniBandThis is a cluster tutorial, not a managed model endpoint

The live table below shows rental comparisons. Use it after selecting the engine and memory requirement, then confirm the deployment image matches your intended version.

GPUCheapest $/GPU-hrProviderProviders in stock
H100$2.59QuantaCloud7
H200$3.43QuantaCloud7
L40S$0.97Massed Compute5
RTX 4090$0.43Vast.ai3
MI300Xnone in stock
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

For a template, review who owns upgrades, restarts and endpoint configuration. For a managed endpoint, ask which engine settings you can change and how you can pin or roll back a version. Keep a known working configuration before changing the image. Choose the deployment whose operating responsibilities your team can actually take on.

The starting decision

These are editorial starting recommendations based on the documented limits above, not measured performance rankings.

SituationEngine to start with
Shared NVIDIA or AMD endpoint, no specialized requirement yetvLLM
Long conversations or repeated prefixes dominateCompare SGLang with the same vLLM workload
Supported NVIDIA model and a need for NVIDIA-specific kernelsTensorRT-LLM, pinned to a specific release
Individual developer wants model managementOllama
CPU/GPU offload, Apple deployment or direct GGUF controlllama.cpp
Existing TGI endpoint needs a future maintenance pathCreate and validate a vLLM or SGLang replacement
Multiple model frameworks or cluster routing requirementsEvaluate Triton or Dynamo around the selected engine

Start with vLLM for the rented shared endpoint. Switch only when a supported feature or a workload-matched comparison gives you a concrete reason.

Sources

Frequently asked questions

Which should I start with, SGLang or vLLM?▾

Start with vLLM for a shared serving endpoint, then compare SGLang when repeated prefixes and long conversations dominate your traffic. This is a starting recommendation, not a speed ranking.

Does TensorRT-LLM still require an engine build?▾

NVIDIA's migration guide, updated 21 September 2026, says the current backend removes separate checkpoint conversion and trtllm-build. Pin your release because older deployment instructions differ.

Is Ollama limited to one user?▾

No. Ollama documents concurrent requests and multi-GPU placement, but its request parallelism and context length increase memory needs.

Is TGI still a good default for a new deployment?▾

No. Hugging Face archived its repository on 21 March 2026 and recommends vLLM or SGLang for endpoints going forward.

Is Triton Inference Server another LLM engine?▾

NVIDIA describes Triton as a general model server with backends, including TensorRT-LLM. Its separate OpenAI frontend supports TensorRT-LLM orchestrator mode, but not leader mode, for model parallelism.

Related Posts