Interactive Hardware Sizing

LLM GPU VRAM & Inference Calculator

Calculate exact GPU memory requirements, KV cache footprint, and hardware suitability for hosting LLMs on vLLM, SGLang, TensorRT-LLM, and Ollama.

1. Select Model Preset
2. Configuration & Precision
Model Parameters: 8.0 Billion
Max Context Window (Tokens): 8,192 tokens
Concurrent Requests (Batch Size): 4 concurrent
Total Estimated VRAM Required
12.4 GB
WEIGHTS
8.0 GB
KV CACHE
2.4 GB
OVERHEAD
2.0 GB
Hardware Compatibility
vLLM Launch Command
vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.90 --max-model-len 8192

How LLM VRAM Calculation Works

1. Model Weight Memory

Calculated as Parameters × Bytes Per Weight. FP16 uses 2 bytes/parameter, FP8 uses 1 byte, and INT4 uses 0.5 bytes. A 70B model in FP16 requires 140GB just to load the weights.

2. KV Cache Footprint

Calculated as 2 × Layers × Heads × Dim × Precision × Context × Batch. Modern architectures with Grouped-Query Attention (GQA) reduce KV cache memory by 4x to 8x compared to Multi-Head Attention.

3. CUDA Context & Buffers

Every PyTorch/CUDA runtime reserves ~1.5GB to 2.5GB for activation tensors, CUDA kernels, and memory allocator fragmentation before generating tokens.