AI Engineering #llm inference#vllm#sglang#tensorrt#gpu optimization#performance

Fast LLM Serving in Production: vLLM, SGLang, and TensorRT-LLM

S

S L Manikanta

Aug 22, 2026 • 11 min read

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

[!NOTE] Quick Engine Selection Rule:

  • Choose vLLM if you need the widest model support, multi-LoRA switching, and standard production stability.
  • Choose SGLang for complex multi-turn prompts, tool-calling agents, and structured JSON generation (RadixAttention gives up to 5x higher throughput on repeated prefixes).
  • Choose TensorRT-LLM for maximum bare-metal GPU performance on static, monolithic enterprise models on NVIDIA H100/A100 hardware.

Quick Inference Engine Comparison

DimensionvLLM (v0.6+)SGLang (v0.4+)TensorRT-LLM
KV Cache ArchitecturePagedAttentionRadixAttention (LRU Prefix Tree)Paged KV with In-Flight Batching
Prefix Caching SpeedLinear lookup / HashTree-based radix cache (Fastest)Optimized tensor engine
Setup ComplexityLow (pip install vllm)Low (pip install sglang)High (Model build/compilation required)
Tool Calling / Structured JSONOutlines / Guided DecodersNative Jump-Forward DecodingCustom C++ plugins
Recommended Use CaseGeneral API Serving & RAGAgentic Loops & Structured OutputMaximum Static Throughput

Running large language models in production gets expensive quickly if you use default settings. Most teams find their GPUs sitting at under 25% compute utilization while user requests suffer from slow response times and sudden latency spikes.

The reason is simple: generating text token by token is limited by GPU memory bandwidth, not raw compute power. The GPU spends most of its time copying model weights and past token keys from GPU memory into the processor.

To fix this and get high throughput, you need three things:

  1. Smart memory management so memory is not wasted on empty space.
  2. Continuous batching so the GPU never sits idle waiting for slow requests to finish.
  3. Quantization to fit more requests into memory at the same time.

Here is a practical breakdown of how the top engines work, how they compare, and how to run them in production.

flowchart TD
    Client[Incoming Client Requests] --> Gateway[API Gateway / Load Balancer]
    Gateway --> Scheduler[Continuous Batch Scheduler]

    subgraph Engine [Inference Server]
        Scheduler --> MemoryManager[KV Cache Manager]
        MemoryManager --> PrefillEngine[Prefill Stage: Reads Full Prompt]
        MemoryManager --> DecodeEngine[Decode Stage: Generates One Token at a Time]
        
        PrefillEngine --> PagedKV[Paged / Radix Memory Pool]
        DecodeEngine --> PagedKV
    end

    PagedKV --> GPU[GPU Memory and Tensor Cores]
    GPU --> Output[Streaming Text Response]

Why LLMs Are Slow: Prefill vs Decode

Every request goes through two separate stages. Understanding the difference is key to tuning your setup.

1. The Prefill Stage (Reading the Prompt)

When a user sends a prompt, the GPU reads all the words in the prompt at the same time. This step uses a lot of compute. The GPU is doing matrix math across all input tokens together.

The metric you track here is Time to First Token (TTFT).

2. The Decode Stage (Generating the Answer)

Once the prompt is read, the model generates words one at a time. To generate a single word, the GPU must load every single model weight and every past token from GPU RAM (High Bandwidth Memory) into the processor cores.

Because very little math happens per byte transferred, the GPU memory speed becomes the hard limit.

The metric you track here is Inter Token Latency (ITL), which is the time between each output token.

Compute per Byte of Memory
▲
│  Prefill (Prompt Processing)
│  ┌───────────────────────────────┐
│  │ Compute Heavy Area            │
│  └───────────────────────────────┘
│
│  Decode (Token Generation)
│  ┌───────────────────────────────┐
│  │ Memory Speed Bottleneck Area  │
│  └───────────────────────────────┘
└────────────────────────────────────────────────────────► Batch Size

When a user sends a huge 8,000 word document while other users are already streaming answers, a basic server will pause everyone else’s streams just to read the big document. That causes massive lag spikes.


Comparing the Top Engines: vLLM, SGLang, and TensorRT-LLM

Three open source engines dominate production setups today. Each is built for a slightly different use case.

FeaturevLLMSGLangTensorRT-LLM
Main Memory TrickPagedAttention (allocates memory in fixed pages)RadixAttention (remembers shared prompt prefixes in a tree)In flight batching with custom C++ memory buffers
Best ForGeneral purpose API servers, easy model setupMulti turn chat, AI agents with tools, JSON outputGetting maximum speed from fixed models on NVIDIA clusters
Prefix CachingChunk based cachingAutomatic Radix tree prefix matchingBlock based KV cache reuse
Chunked PrefillBuilt in (--enable-chunked-prefill)Built inBuilt in
QuantizationFP8, AWQ, GPTQ, BitsAndBytesFP8, AWQ, GPTQ, MarlinFP8, FP4, INT4, INT8 SmoothQuant
Ease of UseVery easy, pure Python with fast C++ kernelsEasy, fast Python and CUDA backendComplex, requires compiling custom model engines

1. vLLM and PagedAttention

In the past, servers reserved memory for the longest possible answer up front. If you configured a 4,000 token limit, the server reserved all 4,000 slots even if the user only asked for a three word reply. This wasted over 60% of GPU memory.

vLLM fixed this with PagedAttention. It breaks GPU memory into small pages (like 16 tokens each) and only allocates a new page when the model actually produces new words. This lets you run 2x to 4x more concurrent users on the same GPU.

2. SGLang and RadixAttention

AI agents and chat bots often send the exact same system prompt and tool definitions with every request.

SGLang keeps track of these shared prompts using a Radix tree in GPU memory. When a new request arrives with a familiar system prompt, SGLang does not re-read the prompt. It reuses the past memory immediately. For agent workflows, this cuts response latency drastically.

3. TensorRT-LLM

This is NVIDIA’s dedicated inference engine. It turns models into compiled binaries with deeply fused CUDA kernels. It gives you the highest raw speed and lowest latency, but you have to compile custom engine files for each specific GPU model.


Three Key Settings to Maximize Speed

1. Continuous Batching

Traditional servers waited for a full batch of 16 requests before starting. If one request finished in 10 tokens and another took 500 tokens, the first slot sat completely empty and wasted for 490 iterations.

Continuous batching checks the batch after every single token. The moment one request finishes, the server immediately inserts a new request from the queue into that open slot.

gantt
    title Static Batching vs Continuous Batching
    dateFormat  X
    axisFormat %s
    
    section Old Static Batching
    Request 1 (50 tokens)       :active, 0, 50
    Request 2 (150 tokens)      :crit, 0, 150
    Request 3 (80 tokens)       :active, 0, 80
    Wasted GPU Idle Time        :milestone, 80, 150
    
    section Continuous Batching
    Request 1 (50 tokens)       :active, 0, 50
    Request 4 (Takes open slot) :done, 50, 110
    Request 2 (150 tokens)      :crit, 0, 150
    Request 3 (80 tokens)       :active, 0, 80
    Request 5 (Takes open slot) :done, 80, 140

2. Chunked Prefill

If someone sends a 10,000 token prompt, processing it all at once can take 400ms. During that time, all other active users experience a freeze in their streaming text.

Chunked prefill splits large prompts into smaller pieces (like 512 tokens). In each step, the GPU processes one small chunk of the big prompt along with one token for all active streams. This keeps streaming smooth for everyone.

3. FP8 KV Cache

The memory needed to store past tokens (the KV cache) grows quickly:

For a 70B model in standard 16 bit precision, storing the cache for 200 users at 4,000 tokens takes over 100 GB of GPU memory.

Switching the KV cache to FP8 (8 bit precision) cuts that memory requirement in half to 50 GB. The drop in model output quality is so small it is almost impossible to measure, but you can now handle double the users on the same hardware.


Ready to Run Production Setups

Running vLLM for High Concurrency

Here is a tested startup script for running a large model across 4 GPUs with FP8 cache and chunked prefill:

#!/usr/bin/env bash
set -euo pipefail

python3 -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 4 \
    --dtype bfloat16 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.92 \
    --max-model-len 16384 \
    --max-num-batched-tokens 8192 \
    --max-num-seqs 256 \
    --enable-chunked-prefill=true \
    --enable-prefix-caching \
    --port 8000 \
    --disable-log-requests

What these flags do:

  • --tensor-parallel-size 4: Splits the model across 4 GPUs.
  • --kv-cache-dtype fp8: Halves the cache memory footprint.
  • --gpu-memory-utilization 0.92: Uses 92% of free GPU memory for token storage.
  • --enable-chunked-prefill=true: Prevents large prompts from freezing existing streams.
  • --enable-prefix-caching: Automatically reuses cached memory for repeated prompt prefixes.

Running SGLang for AI Agents

For agent workloads with lots of tool calls and prompt reuse, use SGLang:

#!/usr/bin/env bash
set -euo pipefail

python3 -m sglang.launch_server \
    --model-path deepseek-ai/DeepSeek-V3 \
    --tp 8 \
    --mem-fraction-static 0.88 \
    --context-length 32768 \
    --quantization fp8 \
    --kv-cache-dtype fp8_e5m2 \
    --schedule-policy lpm \
    --host 0.0.0.0 \
    --port 30000

The --schedule-policy lpm flag tells SGLang to pick requests that match cached prefixes first, maximizing memory reuse.


Testing Your Server Under Real Load

Do not use simple HTTP benchmark tools like curl or ApacheBench to test LLM servers. You need a script that actually streams and measures time to first token and time between tokens.

Here is a Python script using httpx to measure real performance under load:

import asyncio
import time
import statistics
import httpx
from typing import List, Dict

API_URL = "http://localhost:8000/v1/chat/completions"
MODEL = "meta-llama/Llama-3.1-70B-Instruct"

async def test_stream(client: httpx.AsyncClient, prompt: str, max_tokens: int) -> Dict[str, float]:
    payload = {
        "model": MODEL,
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": max_tokens,
        "temperature": 0.0,
        "stream": True,
    }
    
    start_time = time.perf_counter()
    first_token_time = None
    token_times: List[float] = []

    async with client.stream("POST", API_URL, json=payload, timeout=60.0) as response:
        if response.status_code != 200:
            raise RuntimeError(f"Request failed with status {response.status_code}")
            
        async for line in response.aiter_lines():
            if line.startswith("data: ") and line.strip() != "data: [DONE]":
                now = time.perf_counter()
                if first_token_time is None:
                    first_token_time = now
                token_times.append(now)

    end_time = time.perf_counter()
    total_tokens = len(token_times)
    ttft_ms = (first_token_time - start_time) * 1000 if first_token_time else 0.0
    duration = end_time - start_time
    
    # Calculate time between tokens
    itls = [(token_times[i] - token_times[i-1]) * 1000 for i in range(1, len(token_times))]
    p95_itl = statistics.quantiles(itls, n=20)[18] if len(itls) >= 20 else 0.0
    
    return {
        "ttft_ms": ttft_ms,
        "p95_itl_ms": p95_itl,
        "tokens_per_sec": total_tokens / duration if duration > 0 else 0.0,
        "tokens": total_tokens,
    }

async def run_load_test(concurrency: int = 32):
    prompt = "Explain how database indexing works in detail. " * 20
    async with httpx.AsyncClient(limits=httpx.Limits(max_connections=concurrency * 2)) as client:
        tasks = [test_stream(client, prompt, max_tokens=100) for _ in range(concurrency)]
        results = await asyncio.gather(*tasks, return_exceptions=True)
        
    valid = [r for r in results if isinstance(r, dict)]
    avg_ttft = statistics.mean([r["ttft_ms"] for r in valid])
    total_tps = sum([r["tokens_per_sec"] for r in valid])
    
    print(f"Results for {concurrency} concurrent streams:")
    print(f"Success: {len(valid)}/{concurrency}")
    print(f"Average Time to First Token: {avg_ttft:.2f} ms")
    print(f"Total Throughput: {total_tps:.2f} tokens/sec")

if __name__ == "__main__":
    asyncio.run(run_load_test(concurrency=32))

Common Production Problems and Quick Fixes

1. Memory Exhaustion and Stream Stalls

  • What happens: The server suddenly slows to a crawl and token throughput drops.
  • Why: The server accepted more requests than it had memory pages for. It was forced to pause active streams and dump their cache to CPU RAM.
  • Fix: Set --max-num-seqs to a realistic number and keep --gpu-memory-utilization around 0.90 so you have a safety margin.

2. Multi GPU NUMA Slowdown

  • What happens: One GPU is noticeably slower than the others in a multi GPU box.
  • Why: The process running that GPU is talking across CPU sockets over a slow bridge.
  • Fix: Pin your Python processes to the local CPU socket using numactl:
numactl --cpunodebind=0 --membind=0 python3 -m vllm.entrypoints.openai.api_server ...

3. Out of Memory Errors During Long Runs

  • What happens: PyTorch throws CUDA out of memory even though your cache usage shows room.
  • Why: PyTorch memory allocator gets fragmented over time from variable tensor sizes.
  • Fix: Turn on expandable segments in PyTorch before starting your server:
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

Which Engine Should You Pick?

  • Pick vLLM if you want the easiest setup, wide model support, and great stability for standard chat and text generation APIs.
  • Pick SGLang if you are building complex AI agents, tool pipelines, or multi turn chats where reusing prompt memory saves huge amounts of time.
  • Pick TensorRT-LLM if you have a dedicated cluster of NVIDIA GPUs running one specific model at massive scale and you want every last drop of speed.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

AI Engineering
Advanced RAG on Azure: Hybrid Search & Re-ranking Implementation

Going beyond basic vector search. A technical guide to implementing Hybrid Search (Keyword + Vector) and Semantic Re-ranking using Azure AI Search and OpenAI.

AI Engineering
The Shift to Agentic AI: Why Enterprise Architecture is Moving Beyond Chatbots

Chatbots are dead. Welcome to the era of Agentic AI. Explore how enterprises are deploying autonomous agents for complex workflows, the architectural shift required, and the rise of specialized inference models like Nemotron 3.5 Lightning.

AI Engineering
Mastering Agent Skills: A New Standard for AI Capabilities

An in-depth guide on Agent Skills, exploring how to extend AI agents like Claude with specialized knowledge, workflows, and tools using an open, filesystem-based format.