Llama · Llama 4 · 400B parameters · 17B active

Llama 4 Maverick 400B-A17B VRAM requirements

Llama 4 Maverick 400B-A17B has 48 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 226.0 GB.

Won't fit

228.3 GB of 21.8 GB · 1049%
0237 GB
Weights 226.0 GB
KV cache 1.5 GB
Runtime overhead 0.8 GB
Over the limit 206.6 GB

Short by 206.6 GB. You can run it with 4 of 48 layers on the RTX 4090 and the rest in system RAM, at roughly 4.00 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation4.00tok/s
Prompt processing2038tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Llama 4 Maverick 400B-A17B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 745.1 GB 747.4 GB Won't fit — 1.20
INT8 / W8A8 8.50 395.3 GB 397.7 GB Won't fit — 2.27
Q8_0 (GGUF) 8.50 395.3 GB 397.7 GB Won't fit — 2.27
FP8 (E4M3) 8.00 372.5 GB 374.9 GB Won't fit — 2.40
Q6_K 6.56 305.5 GB 307.8 GB Won't fit — 2.95
Q5_K_M 5.67 264.0 GB 266.4 GB Won't fit — 3.39
Q5_K_S 5.52 257.0 GB 259.4 GB Won't fit — 3.47
Q4_K_M 4.85 226.0 GB 228.3 GB Won't fit — 4.00
Q4_K_S 4.58 213.5 GB 215.8 GB Won't fit — 4.21
Q4_0 4.55 212.1 GB 214.4 GB Won't fit — 4.24
AWQ 4-bit 4.25 200.7 GB 203.1 GB Won't fit — 4.46
GPTQ 4-bit 4.25 200.7 GB 203.1 GB Won't fit — 4.46
MXFP4 4.25 200.7 GB 203.1 GB Won't fit — 4.46
IQ4_XS 4.25 198.2 GB 200.5 GB Won't fit — 4.51
Q3_K_M 3.91 182.5 GB 184.8 GB Won't fit — 4.97
IQ3_M 3.70 172.7 GB 175.1 GB Won't fit — 5.22
IQ3_XXS 3.06 143.1 GB 145.4 GB Won't fit — 6.32
Q2_K 2.63 123.2 GB 125.5 GB Won't fit — 7.37
IQ2_XXS 2.06 96.8 GB 99.1 GB Won't fit — 9.48
IQ1_M 1.75 82.4 GB 84.7 GB Won't fit — 11.4

Llama 4 Maverick 400B-A17B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
Mac Studio M3 Ultra 256GB 256 819 Won't fit — 14.6
H100 SXM 80GB 80 3350 Won't fit — 6.20
Mac Studio M4 Max 128GB 128 546 Won't fit — 6.01
NVIDIA DGX Spark (GB10) 128 273 Won't fit — 5.71
A100 80GB 80 2039 Won't fit — 5.51
RTX A6000 48 768 Won't fit — 4.53
RTX 5090 32 1792 Won't fit — 4.52
Ryzen AI Max+ 395 128GB 128 256 Won't fit — 4.49
Mac Mini M4 Pro 48GB 48 273 Won't fit — 4.43
L40S 48 864 Won't fit — 4.35
RTX 5080 16 960 Won't fit — 4.23
RTX 5070 Ti 16 896 Won't fit — 4.23
RTX 5060 Ti 16GB 16 448 Won't fit — 4.21
RTX 3090 24 936 Won't fit — 4.17
RTX 5070 12 672 Won't fit — 4.14
RTX 4090 24 1008 Won't fit — 4.00
RTX 3080 10GB 10 760 Won't fit — 3.92
RTX 3060 12GB 12 360 Won't fit — 3.92
RTX 4080 Super 16 736 Won't fit — 3.83
RTX 4070 Ti Super 16 672 Won't fit — 3.83
RTX 4060 Ti 16GB 16 288 Won't fit — 3.81
Radeon RX 7900 XTX 24 960 Won't fit — 3.79
RTX 4070 Super 12 504 Won't fit — 3.76
RTX 4070 12 504 Won't fit — 3.76
Arc B580 12 456 Won't fit — 3.18

Architecture

Parameters400B
Active per token17B of 128 experts, top-1
Layers48
Hidden size5120
Attention heads / KV heads40 / 8
Head dimension128
Vocabulary202,048
Trained context1024K
KV cache per 1K tokens0 GB
Hugging Facemeta-llama/Llama-4-Maverick-17B-128E-Instruct

The Llama family

Meta's open-weight series, and the default target for most local tooling. Every Llama 3.x model uses grouped-query attention with 8 KV heads, so the cache stays modest even at 70B. Llama 4 moved to mixture-of-experts: Scout and Maverick occupy 109B and 400B of memory but read only 17B per token.

huggingface.co/meta-llama · llama.com · all 7 Llama models

Direct answers