Llama · Llama 3.1 · 405.9B parameters

Llama 3.1 405B VRAM requirements

Llama 3.1 405B has 126 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 229.5 GB.

Won't fit

234.4 GB of 21.8 GB · 1077%
0244 GB
Weights 229.5 GB
KV cache 3.9 GB
Runtime overhead 1.0 GB
Over the limit 212.7 GB

Short by 212.7 GB. You can run it with 9 of 126 layers on the RTX 4090 and the rest in system RAM, at roughly 0.18 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation0.18tok/s
Prompt processing85.4tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 405B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 756.0 GB 760.9 GB Won't fit — 0.05
INT8 / W8A8 8.50 400.7 GB 405.6 GB Won't fit — 0.10
Q8_0 (GGUF) 8.50 400.7 GB 405.6 GB Won't fit — 0.10
FP8 (E4M3) 8.00 378.0 GB 382.9 GB Won't fit — 0.10
Q6_K 6.56 309.9 GB 314.9 GB Won't fit — 0.13
Q5_K_M 5.67 267.9 GB 272.9 GB Won't fit — 0.15
Q5_K_S 5.52 260.8 GB 265.8 GB Won't fit — 0.15
Q4_K_M 4.85 229.5 GB 234.4 GB Won't fit — 0.18
Q4_K_S 4.58 216.8 GB 221.8 GB Won't fit — 0.19
Q4_0 4.55 215.4 GB 220.4 GB Won't fit — 0.19
AWQ 4-bit 4.25 206.5 GB 211.5 GB Won't fit — 0.20
GPTQ 4-bit 4.25 206.5 GB 211.5 GB Won't fit — 0.20
MXFP4 4.25 206.5 GB 211.5 GB Won't fit — 0.20
IQ4_XS 4.25 201.4 GB 206.4 GB Won't fit — 0.20
Q3_K_M 3.91 185.5 GB 190.5 GB Won't fit — 0.22
IQ3_M 3.70 175.7 GB 180.7 GB Won't fit — 0.24
IQ3_XXS 3.06 145.8 GB 150.7 GB Won't fit — 0.29
Q2_K 2.63 125.7 GB 130.6 GB Won't fit — 0.34
IQ2_XXS 2.06 99.0 GB 104.0 GB Won't fit — 0.45
IQ1_M 1.75 84.5 GB 89.5 GB Won't fit — 0.54

Llama 3.1 405B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
Mac Studio M3 Ultra 256GB 256 819 Won't fit — 0.63
H100 SXM 80GB 80 3350 Won't fit — 0.27
Mac Studio M4 Max 128GB 128 546 Won't fit — 0.27
NVIDIA DGX Spark (GB10) 128 273 Won't fit — 0.26
A100 80GB 80 2039 Won't fit — 0.24
Ryzen AI Max+ 395 128GB 128 256 Won't fit — 0.21
RTX A6000 48 768 Won't fit — 0.20
RTX 5090 32 1792 Won't fit — 0.20
Mac Mini M4 Pro 48GB 48 273 Won't fit — 0.20
L40S 48 864 Won't fit — 0.20
RTX 5080 16 960 Won't fit — 0.19
RTX 5070 Ti 16 896 Won't fit — 0.19
RTX 5060 Ti 16GB 16 448 Won't fit — 0.19
RTX 5070 12 672 Won't fit — 0.19
RTX 3090 24 936 Won't fit — 0.18
RTX 4090 24 1008 Won't fit — 0.18
RTX 3060 12GB 12 360 Won't fit — 0.18
RTX 3080 10GB 10 760 Won't fit — 0.17
RTX 4080 Super 16 736 Won't fit — 0.17
RTX 4070 Ti Super 16 672 Won't fit — 0.17
RTX 4060 Ti 16GB 16 288 Won't fit — 0.17
RTX 4070 Super 12 504 Won't fit — 0.17
RTX 4070 12 504 Won't fit — 0.17
Radeon RX 7900 XTX 24 960 Won't fit — 0.17
Arc B580 12 456 Won't fit — 0.14

Architecture

Parameters405.9B
Layers126
Hidden size16384
Attention heads / KV heads128 / 8
Head dimension128
Vocabulary128,256
Trained context128K
KV cache per 1K tokens0 GB
Hugging Facemeta-llama/Llama-3.1-405B-Instruct

The Llama family

Meta's open-weight series, and the default target for most local tooling. Every Llama 3.x model uses grouped-query attention with 8 KV heads, so the cache stays modest even at 70B. Llama 4 moved to mixture-of-experts: Scout and Maverick occupy 109B and 400B of memory but read only 17B per token.

huggingface.co/meta-llama · llama.com · all 7 Llama models

Direct answers