AI Engineering #deepseek r1#local llm#ollama#vllm#hardware#reasoning models

How to Run DeepSeek R1 Locally: Hardware Requirements and Setup Guide

S

S L Manikanta

Aug 27, 2026 • 5 min read

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

DeepSeek R1 demonstrated that open weights reasoning models can match proprietary frontier models like OpenAI o1 on math, coding, and logical tasks. Because the model weights are open, developers and teams can run DeepSeek R1 completely locally on their own hardware.

However, deciding which version to run depends heavily on your available GPU VRAM and system memory. DeepSeek released both distilled models (ranging from 1.5B to 70B parameters) and the full 671B parameter Mixture of Experts (MoE) model.

Here is a breakdown of the hardware requirements for every model size, along with setup instructions using Ollama for single machines and vLLM for multi-GPU servers.

flowchart TD
    User[Target Hardware Environment] --> Decision{Select Available Hardware}

    Decision -->|Laptop / Mac Mini / 8GB-16GB RAM| DistillSmall[Run 1.5B, 7B, 8B Distills via Ollama]
    Decision -->|Workstation / Single 24GB GPU / 32GB-64GB RAM| DistillMedium[Run 14B, 32B, 70B Distills via Ollama / llama.cpp]
    Decision -->|Enterprise Server / Multi GPU 8x H100 / A100| FullModel[Run Full 671B MoE via vLLM / SGLang FP8]

    DistillSmall --> PrivateLocal[Private Local Chat & Code Completion]
    DistillMedium --> LocalCoder[Full Local Agent & Reasoning Workflows]
    FullModel --> EnterpriseAPI[High Throughput Enterprise AI Endpoint]

Hardware Requirements Matrix

Here is the exact VRAM and system RAM you need for each variant of DeepSeek R1:

Model VariantBase ArchitecturePrecision / QuantizationMinimum VRAM / RAMRecommended Hardware
DeepSeek R1 Distill 1.5BQwen 2.5 1.5BQ4_K_M (GGUF)2 GBAny modern laptop, Apple M1/M2/M3 (8GB RAM)
DeepSeek R1 Distill 7B / 8BQwen 2.5 / Llama 3.1Q4_K_M (GGUF)6 GBRTX 3060, RTX 4060, Mac 16GB
DeepSeek R1 Distill 14BQwen 2.5 14BQ4_K_M (GGUF)10 GBRTX 3080, RTX 4070 (12GB), Mac 18GB+
DeepSeek R1 Distill 32BQwen 2.5 32BQ4_K_M (GGUF)20 GBRTX 3090, RTX 4090 (24GB), Mac 36GB+
DeepSeek R1 Distill 70BLlama 3.3 70BQ4_K_M (GGUF)42 GB2x RTX 3090 / 4090, Mac Studio 64GB+
DeepSeek R1 Full (671B MoE)671B Total (37B active)FP8 (Compressed)360 GB8x H100 (80GB), 8x A100 (80GB)
DeepSeek R1 Full (671B MoE)671B Total (37B active)Q4_K_M (GGUF CPU offload)400 GB RAMDual AMD EPYC / Threadripper with 512GB RAM

Method 1: Running with Ollama (Fastest Local Setup)

Ollama is the easiest way to download and run DeepSeek R1 on macOS, Windows, or Linux. It automatically chooses GPU offloading and sets up an OpenAI compatible API on localhost:11434.

1. Download and Run

Open your terminal and run the model size that fits your VRAM:

# For laptops with 16GB RAM (Excellent daily reasoning assistant)
ollama run deepseek-r1:8b

# For workstations with 24GB VRAM (Sweet spot for coding and complex logic)
ollama run deepseek-r1:14b
ollama run deepseek-r1:32b

# For Mac Studio or dual GPU setups
ollama run deepseek-r1:70b

2. Using the OpenAI Compatible API

Once running, you can connect your existing Python scripts, LangChain agents, or IDE extensions (like Cursor or Continue) directly:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama", # Required by SDK but unused locally
)

response = client.chat.completions.create(
    model="deepseek-r1:14b",
    messages=[
        {"role": "user", "content": "Write a Python script that detects cycles in a directed graph using Kahn's algorithm."}
    ],
    temperature=0.6,
)

# DeepSeek R1 outputs internal reasoning inside <think> tags before the answer
print(response.choices[0].message.content)

Method 2: Serving DeepSeek R1 with vLLM (Production Multi GPU)

If you have a dedicated server with multiple GPUs and want high throughput token generation with continuous batching, use vLLM.

1. Install vLLM

pip install vllm

2. Launching Distill 70B on 2x RTX 4090 or 2x A10 GPUs

python3 -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-R1-Distill-Llama-70B \
    --tensor-parallel-size 2 \
    --dtype bfloat16 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.90 \
    --max-model-len 16384 \
    --port 8000

3. Launching the Full 671B MoE on an 8x H100 Cluster

To serve the full 671B model at scale, use FP8 precision across 8 GPUs:

python3 -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-R1 \
    --tensor-parallel-size 8 \
    --quantization fp8 \
    --kv-cache-dtype fp8 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.94 \
    --port 8000

Best Practices for Prompting DeepSeek R1

Reasoning models behave differently from standard chat models:

  1. Avoid Strict System Prompts: DeepSeek R1 performs best with a clean user prompt. Overly complex system instructions can interfere with its internal thinking chain.
  2. Temperature Settings: Keep the temperature between 0.5 and 0.7. Setting temperature to 0.0 can cause repetitive reasoning loops, while values above 0.8 can cause hallucinations.
  3. Handling the <think> Block: In UI applications, parse and collapse the <think>...</think> section so users can expand the reasoning chain if needed without cluttering the main response.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

AI Engineering
Advanced RAG on Azure: Hybrid Search & Re-ranking Implementation

Going beyond basic vector search. A technical guide to implementing Hybrid Search (Keyword + Vector) and Semantic Re-ranking using Azure AI Search and OpenAI.

AI Engineering
The Shift to Agentic AI: Why Enterprise Architecture is Moving Beyond Chatbots

Chatbots are dead. Welcome to the era of Agentic AI. Explore how enterprises are deploying autonomous agents for complex workflows, the architectural shift required, and the rise of specialized inference models like Nemotron 3.5 Lightning.

AI Engineering
Mastering Agent Skills: A New Standard for AI Capabilities

An in-depth guide on Agent Skills, exploring how to extend AI agents like Claude with specialized knowledge, workflows, and tools using an open, filesystem-based format.