Fast LLM Serving in Production: vLLM, SGLang, and TensorRT-LLM
S L Manikanta
Aug 22, 2026 • 11 min read
list On this page expand_more
- Quick Inference Engine Comparison
- Why LLMs Are Slow: Prefill vs Decode
- 1. The Prefill Stage (Reading the Prompt)
- 2. The Decode Stage (Generating the Answer)
- Comparing the Top Engines: vLLM, SGLang, and TensorRT-LLM
- 1. vLLM and PagedAttention
- 2. SGLang and RadixAttention
- 3. TensorRT-LLM
- Three Key Settings to Maximize Speed
- 1. Continuous Batching
- 2. Chunked Prefill
- 3. FP8 KV Cache
- Ready to Run Production Setups
- Running vLLM for High Concurrency
- Running SGLang for AI Agents
- Testing Your Server Under Real Load
- Common Production Problems and Quick Fixes
- 1. Memory Exhaustion and Stream Stalls
- 2. Multi GPU NUMA Slowdown
- 3. Out of Memory Errors During Long Runs
- Which Engine Should You Pick?
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
[!NOTE] Quick Engine Selection Rule:
- Choose vLLM if you need the widest model support, multi-LoRA switching, and standard production stability.
- Choose SGLang for complex multi-turn prompts, tool-calling agents, and structured JSON generation (RadixAttention gives up to 5x higher throughput on repeated prefixes).
- Choose TensorRT-LLM for maximum bare-metal GPU performance on static, monolithic enterprise models on NVIDIA H100/A100 hardware.
Quick Inference Engine Comparison
| Dimension | vLLM (v0.6+) | SGLang (v0.4+) | TensorRT-LLM |
|---|---|---|---|
| KV Cache Architecture | PagedAttention | RadixAttention (LRU Prefix Tree) | Paged KV with In-Flight Batching |
| Prefix Caching Speed | Linear lookup / Hash | Tree-based radix cache (Fastest) | Optimized tensor engine |
| Setup Complexity | Low (pip install vllm) | Low (pip install sglang) | High (Model build/compilation required) |
| Tool Calling / Structured JSON | Outlines / Guided Decoders | Native Jump-Forward Decoding | Custom C++ plugins |
| Recommended Use Case | General API Serving & RAG | Agentic Loops & Structured Output | Maximum Static Throughput |
Running large language models in production gets expensive quickly if you use default settings. Most teams find their GPUs sitting at under 25% compute utilization while user requests suffer from slow response times and sudden latency spikes.
The reason is simple: generating text token by token is limited by GPU memory bandwidth, not raw compute power. The GPU spends most of its time copying model weights and past token keys from GPU memory into the processor.
To fix this and get high throughput, you need three things:
- Smart memory management so memory is not wasted on empty space.
- Continuous batching so the GPU never sits idle waiting for slow requests to finish.
- Quantization to fit more requests into memory at the same time.
Here is a practical breakdown of how the top engines work, how they compare, and how to run them in production.
flowchart TD
Client[Incoming Client Requests] --> Gateway[API Gateway / Load Balancer]
Gateway --> Scheduler[Continuous Batch Scheduler]
subgraph Engine [Inference Server]
Scheduler --> MemoryManager[KV Cache Manager]
MemoryManager --> PrefillEngine[Prefill Stage: Reads Full Prompt]
MemoryManager --> DecodeEngine[Decode Stage: Generates One Token at a Time]
PrefillEngine --> PagedKV[Paged / Radix Memory Pool]
DecodeEngine --> PagedKV
end
PagedKV --> GPU[GPU Memory and Tensor Cores]
GPU --> Output[Streaming Text Response]
Why LLMs Are Slow: Prefill vs Decode
Every request goes through two separate stages. Understanding the difference is key to tuning your setup.
1. The Prefill Stage (Reading the Prompt)
When a user sends a prompt, the GPU reads all the words in the prompt at the same time. This step uses a lot of compute. The GPU is doing matrix math across all input tokens together.
The metric you track here is Time to First Token (TTFT).
2. The Decode Stage (Generating the Answer)
Once the prompt is read, the model generates words one at a time. To generate a single word, the GPU must load every single model weight and every past token from GPU RAM (High Bandwidth Memory) into the processor cores.
Because very little math happens per byte transferred, the GPU memory speed becomes the hard limit.
The metric you track here is Inter Token Latency (ITL), which is the time between each output token.
Compute per Byte of Memory
▲
│ Prefill (Prompt Processing)
│ ┌───────────────────────────────┐
│ │ Compute Heavy Area │
│ └───────────────────────────────┘
│
│ Decode (Token Generation)
│ ┌───────────────────────────────┐
│ │ Memory Speed Bottleneck Area │
│ └───────────────────────────────┘
└────────────────────────────────────────────────────────► Batch Size
When a user sends a huge 8,000 word document while other users are already streaming answers, a basic server will pause everyone else’s streams just to read the big document. That causes massive lag spikes.
Comparing the Top Engines: vLLM, SGLang, and TensorRT-LLM
Three open source engines dominate production setups today. Each is built for a slightly different use case.
| Feature | vLLM | SGLang | TensorRT-LLM |
|---|---|---|---|
| Main Memory Trick | PagedAttention (allocates memory in fixed pages) | RadixAttention (remembers shared prompt prefixes in a tree) | In flight batching with custom C++ memory buffers |
| Best For | General purpose API servers, easy model setup | Multi turn chat, AI agents with tools, JSON output | Getting maximum speed from fixed models on NVIDIA clusters |
| Prefix Caching | Chunk based caching | Automatic Radix tree prefix matching | Block based KV cache reuse |
| Chunked Prefill | Built in (--enable-chunked-prefill) | Built in | Built in |
| Quantization | FP8, AWQ, GPTQ, BitsAndBytes | FP8, AWQ, GPTQ, Marlin | FP8, FP4, INT4, INT8 SmoothQuant |
| Ease of Use | Very easy, pure Python with fast C++ kernels | Easy, fast Python and CUDA backend | Complex, requires compiling custom model engines |
1. vLLM and PagedAttention
In the past, servers reserved memory for the longest possible answer up front. If you configured a 4,000 token limit, the server reserved all 4,000 slots even if the user only asked for a three word reply. This wasted over 60% of GPU memory.
vLLM fixed this with PagedAttention. It breaks GPU memory into small pages (like 16 tokens each) and only allocates a new page when the model actually produces new words. This lets you run 2x to 4x more concurrent users on the same GPU.
2. SGLang and RadixAttention
AI agents and chat bots often send the exact same system prompt and tool definitions with every request.
SGLang keeps track of these shared prompts using a Radix tree in GPU memory. When a new request arrives with a familiar system prompt, SGLang does not re-read the prompt. It reuses the past memory immediately. For agent workflows, this cuts response latency drastically.
3. TensorRT-LLM
This is NVIDIA’s dedicated inference engine. It turns models into compiled binaries with deeply fused CUDA kernels. It gives you the highest raw speed and lowest latency, but you have to compile custom engine files for each specific GPU model.
Three Key Settings to Maximize Speed
1. Continuous Batching
Traditional servers waited for a full batch of 16 requests before starting. If one request finished in 10 tokens and another took 500 tokens, the first slot sat completely empty and wasted for 490 iterations.
Continuous batching checks the batch after every single token. The moment one request finishes, the server immediately inserts a new request from the queue into that open slot.
gantt
title Static Batching vs Continuous Batching
dateFormat X
axisFormat %s
section Old Static Batching
Request 1 (50 tokens) :active, 0, 50
Request 2 (150 tokens) :crit, 0, 150
Request 3 (80 tokens) :active, 0, 80
Wasted GPU Idle Time :milestone, 80, 150
section Continuous Batching
Request 1 (50 tokens) :active, 0, 50
Request 4 (Takes open slot) :done, 50, 110
Request 2 (150 tokens) :crit, 0, 150
Request 3 (80 tokens) :active, 0, 80
Request 5 (Takes open slot) :done, 80, 140
2. Chunked Prefill
If someone sends a 10,000 token prompt, processing it all at once can take 400ms. During that time, all other active users experience a freeze in their streaming text.
Chunked prefill splits large prompts into smaller pieces (like 512 tokens). In each step, the GPU processes one small chunk of the big prompt along with one token for all active streams. This keeps streaming smooth for everyone.
3. FP8 KV Cache
The memory needed to store past tokens (the KV cache) grows quickly:
For a 70B model in standard 16 bit precision, storing the cache for 200 users at 4,000 tokens takes over 100 GB of GPU memory.
Switching the KV cache to FP8 (8 bit precision) cuts that memory requirement in half to 50 GB. The drop in model output quality is so small it is almost impossible to measure, but you can now handle double the users on the same hardware.
Ready to Run Production Setups
Running vLLM for High Concurrency
Here is a tested startup script for running a large model across 4 GPUs with FP8 cache and chunked prefill:
#!/usr/bin/env bash
set -euo pipefail
python3 -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--dtype bfloat16 \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.92 \
--max-model-len 16384 \
--max-num-batched-tokens 8192 \
--max-num-seqs 256 \
--enable-chunked-prefill=true \
--enable-prefix-caching \
--port 8000 \
--disable-log-requests
What these flags do:
--tensor-parallel-size 4: Splits the model across 4 GPUs.--kv-cache-dtype fp8: Halves the cache memory footprint.--gpu-memory-utilization 0.92: Uses 92% of free GPU memory for token storage.--enable-chunked-prefill=true: Prevents large prompts from freezing existing streams.--enable-prefix-caching: Automatically reuses cached memory for repeated prompt prefixes.
Running SGLang for AI Agents
For agent workloads with lots of tool calls and prompt reuse, use SGLang:
#!/usr/bin/env bash
set -euo pipefail
python3 -m sglang.launch_server \
--model-path deepseek-ai/DeepSeek-V3 \
--tp 8 \
--mem-fraction-static 0.88 \
--context-length 32768 \
--quantization fp8 \
--kv-cache-dtype fp8_e5m2 \
--schedule-policy lpm \
--host 0.0.0.0 \
--port 30000
The --schedule-policy lpm flag tells SGLang to pick requests that match cached prefixes first, maximizing memory reuse.
Testing Your Server Under Real Load
Do not use simple HTTP benchmark tools like curl or ApacheBench to test LLM servers. You need a script that actually streams and measures time to first token and time between tokens.
Here is a Python script using httpx to measure real performance under load:
import asyncio
import time
import statistics
import httpx
from typing import List, Dict
API_URL = "http://localhost:8000/v1/chat/completions"
MODEL = "meta-llama/Llama-3.1-70B-Instruct"
async def test_stream(client: httpx.AsyncClient, prompt: str, max_tokens: int) -> Dict[str, float]:
payload = {
"model": MODEL,
"messages": [{"role": "user", "content": prompt}],
"max_tokens": max_tokens,
"temperature": 0.0,
"stream": True,
}
start_time = time.perf_counter()
first_token_time = None
token_times: List[float] = []
async with client.stream("POST", API_URL, json=payload, timeout=60.0) as response:
if response.status_code != 200:
raise RuntimeError(f"Request failed with status {response.status_code}")
async for line in response.aiter_lines():
if line.startswith("data: ") and line.strip() != "data: [DONE]":
now = time.perf_counter()
if first_token_time is None:
first_token_time = now
token_times.append(now)
end_time = time.perf_counter()
total_tokens = len(token_times)
ttft_ms = (first_token_time - start_time) * 1000 if first_token_time else 0.0
duration = end_time - start_time
# Calculate time between tokens
itls = [(token_times[i] - token_times[i-1]) * 1000 for i in range(1, len(token_times))]
p95_itl = statistics.quantiles(itls, n=20)[18] if len(itls) >= 20 else 0.0
return {
"ttft_ms": ttft_ms,
"p95_itl_ms": p95_itl,
"tokens_per_sec": total_tokens / duration if duration > 0 else 0.0,
"tokens": total_tokens,
}
async def run_load_test(concurrency: int = 32):
prompt = "Explain how database indexing works in detail. " * 20
async with httpx.AsyncClient(limits=httpx.Limits(max_connections=concurrency * 2)) as client:
tasks = [test_stream(client, prompt, max_tokens=100) for _ in range(concurrency)]
results = await asyncio.gather(*tasks, return_exceptions=True)
valid = [r for r in results if isinstance(r, dict)]
avg_ttft = statistics.mean([r["ttft_ms"] for r in valid])
total_tps = sum([r["tokens_per_sec"] for r in valid])
print(f"Results for {concurrency} concurrent streams:")
print(f"Success: {len(valid)}/{concurrency}")
print(f"Average Time to First Token: {avg_ttft:.2f} ms")
print(f"Total Throughput: {total_tps:.2f} tokens/sec")
if __name__ == "__main__":
asyncio.run(run_load_test(concurrency=32))
Common Production Problems and Quick Fixes
1. Memory Exhaustion and Stream Stalls
- What happens: The server suddenly slows to a crawl and token throughput drops.
- Why: The server accepted more requests than it had memory pages for. It was forced to pause active streams and dump their cache to CPU RAM.
- Fix: Set
--max-num-seqsto a realistic number and keep--gpu-memory-utilizationaround 0.90 so you have a safety margin.
2. Multi GPU NUMA Slowdown
- What happens: One GPU is noticeably slower than the others in a multi GPU box.
- Why: The process running that GPU is talking across CPU sockets over a slow bridge.
- Fix: Pin your Python processes to the local CPU socket using
numactl:
numactl --cpunodebind=0 --membind=0 python3 -m vllm.entrypoints.openai.api_server ...
3. Out of Memory Errors During Long Runs
- What happens: PyTorch throws
CUDA out of memoryeven though your cache usage shows room. - Why: PyTorch memory allocator gets fragmented over time from variable tensor sizes.
- Fix: Turn on expandable segments in PyTorch before starting your server:
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
Which Engine Should You Pick?
- Pick vLLM if you want the easiest setup, wide model support, and great stability for standard chat and text generation APIs.
- Pick SGLang if you are building complex AI agents, tool pipelines, or multi turn chats where reusing prompt memory saves huge amounts of time.
- Pick TensorRT-LLM if you have a dedicated cluster of NVIDIA GPUs running one specific model at massive scale and you want every last drop of speed.
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
Written by S L Manikanta
AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.
Related Articles
Advanced RAG on Azure: Hybrid Search & Re-ranking Implementation
Going beyond basic vector search. A technical guide to implementing Hybrid Search (Keyword + Vector) and Semantic Re-ranking using Azure AI Search and OpenAI.
The Shift to Agentic AI: Why Enterprise Architecture is Moving Beyond Chatbots
Chatbots are dead. Welcome to the era of Agentic AI. Explore how enterprises are deploying autonomous agents for complex workflows, the architectural shift required, and the rise of specialized inference models like Nemotron 3.5 Lightning.
Mastering Agent Skills: A New Standard for AI Capabilities
An in-depth guide on Agent Skills, exploring how to extend AI agents like Claude with specialized knowledge, workflows, and tools using an open, filesystem-based format.