Calculate exact GPU memory requirements, KV cache footprint, and hardware suitability for hosting LLMs on vLLM, SGLang, TensorRT-LLM, and Ollama.
vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.90 --max-model-len 8192
Calculated as Parameters × Bytes Per Weight. FP16 uses 2 bytes/parameter, FP8 uses 1 byte, and INT4 uses 0.5 bytes. A 70B model in FP16 requires 140GB just to load the weights.
Calculated as 2 × Layers × Heads × Dim × Precision × Context × Batch. Modern architectures with Grouped-Query Attention (GQA) reduce KV cache memory by 4x to 8x compared to Multi-Head Attention.
Every PyTorch/CUDA runtime reserves ~1.5GB to 2.5GB for activation tensors, CUDA kernels, and memory allocator fragmentation before generating tokens.