DeepSeek · DeepSeek-R1 · 671B parameters · 37B active

DeepSeek-R1 671B-A37B VRAM requirements

DeepSeek-R1 671B-A37B has 61 layers and uses multi-head latent attention (MLA). At Q4_K_M the weights come to 379.0 GB.

Won't fit

380.4 GB of 21.8 GB · 1748%
0396 GB
Weights 379.0 GB
KV cache 0.5 GB
Runtime overhead 0.9 GB
Over the limit 358.6 GB

Short by 358.6 GB. You can run it with 3 of 61 layers on the RTX 4090 and the rest in system RAM, at roughly 1.88 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation1.88tok/s
Prompt processing936tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of DeepSeek-R1 671B-A37B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 1249.8 GB 1251.2 GB Won't fit — 0.55
INT8 / W8A8 8.50 663.6 GB 665.0 GB Won't fit — 1.05
Q8_0 (GGUF) 8.50 663.6 GB 665.0 GB Won't fit — 1.05
FP8 (E4M3) 8.00 624.9 GB 626.3 GB Won't fit — 1.11
Q6_K 6.56 512.4 GB 513.8 GB Won't fit — 1.37
Q5_K_M 5.67 442.9 GB 444.3 GB Won't fit — 1.59
Q5_K_S 5.52 431.2 GB 432.6 GB Won't fit — 1.63
Q4_K_M 4.85 379.0 GB 380.4 GB Won't fit — 1.88
Q4_K_S 4.58 358.0 GB 359.4 GB Won't fit — 1.98
Q4_0 4.55 355.6 GB 357.0 GB Won't fit — 2.00
AWQ 4-bit 4.25 334.5 GB 335.9 GB Won't fit — 2.12
GPTQ 4-bit 4.25 334.5 GB 335.9 GB Won't fit — 2.12
MXFP4 4.25 334.5 GB 335.9 GB Won't fit — 2.12
IQ4_XS 4.25 332.3 GB 333.7 GB Won't fit — 2.13
Q3_K_M 3.91 305.8 GB 307.2 GB Won't fit — 2.35
IQ3_M 3.70 289.4 GB 290.8 GB Won't fit — 2.48
IQ3_XXS 3.06 239.6 GB 241.0 GB Won't fit — 3.03
Q2_K 2.63 206.1 GB 207.5 GB Won't fit — 3.55
IQ2_XXS 2.06 161.7 GB 163.1 GB Won't fit — 4.55
IQ1_M 1.75 137.5 GB 138.9 GB Won't fit — 5.49

DeepSeek-R1 671B-A37B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
Mac Studio M3 Ultra 256GB 256 819 Won't fit — 3.10
H100 SXM 80GB 80 3350 Won't fit — 2.53
Mac Studio M4 Max 128GB 128 546 Won't fit — 2.43
NVIDIA DGX Spark (GB10) 128 273 Won't fit — 2.40
A100 80GB 80 2039 Won't fit — 2.26
RTX 5090 32 1792 Won't fit — 2.10
RTX A6000 48 768 Won't fit — 2.05
Mac Mini M4 Pro 48GB 48 273 Won't fit — 2.04
RTX 5080 16 960 Won't fit — 2.03
RTX 5070 Ti 16 896 Won't fit — 2.03
RTX 5060 Ti 16GB 16 448 Won't fit — 2.03
RTX 5070 12 672 Won't fit — 2.00
L40S 48 864 Won't fit — 1.97
RTX 3090 24 936 Won't fit — 1.96
Ryzen AI Max+ 395 128GB 128 256 Won't fit — 1.90
RTX 3080 10GB 10 760 Won't fit — 1.90
RTX 3060 12GB 12 360 Won't fit — 1.89
RTX 4090 24 1008 Won't fit — 1.88
RTX 4080 Super 16 736 Won't fit — 1.85
RTX 4070 Ti Super 16 672 Won't fit — 1.84
RTX 4060 Ti 16GB 16 288 Won't fit — 1.84
RTX 4070 Super 12 504 Won't fit — 1.82
RTX 4070 12 504 Won't fit — 1.82
Radeon RX 7900 XTX 24 960 Won't fit — 1.78
Arc B580 12 456 Won't fit — 1.54

Architecture

Parameters671B
Active per token37B of 256 experts, top-8
Layers61
Hidden size7168
Attention heads / KV heads128 / 128
Head dimension128
Vocabulary129,280
Trained context160K
KV cache per 1K tokens0 GB
Hugging Facedeepseek-ai/DeepSeek-R1

The DeepSeek family

V3 and R1 share one 671B mixture-of-experts architecture with 37B active per token, and both use multi-head latent attention — a single compressed KV vector per layer instead of a full set of heads. That gives R1 roughly a fifth of the cache per token of a dense 70B. The R1 distills are ordinary Qwen and Llama models fine-tuned on R1 output, and they fit on one card.

huggingface.co/deepseek-ai · deepseek.com · all 5 DeepSeek models

Direct answers