Mistral · 12.3B parameters

Mistral NeMo 12B VRAM requirements

Mistral NeMo 12B has 40 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 7.0 GB, and the best quantisation that fits a 24 GB card is INT8 / W8A8.

Runs comfortably

9.1 GB of 21.8 GB · 42%
022 GB
Weights 7.0 GB
KV cache 1.3 GB
Runtime overhead 0.8 GB

Mistral NeMo 12B at Q4_K_M leaves 12.7 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation77.6tok/s
Prompt processing2829tok/s
Max context89Ktokens
KV per 1K tokens0GB

Every quantisation of Mistral NeMo 12B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 22.8 GB 24.9 GB Won't fit 8.00
INT8 / W8A8 8.50 11.8 GB 13.9 GB Runs comfortably 58K 48.0
Q8_0 (GGUF) 8.50 11.8 GB 13.9 GB Runs comfortably 58K 48.0
FP8 (E4M3) 8.00 11.4 GB 13.5 GB Runs comfortably 61K 49.6
Q6_K 6.56 9.4 GB 11.4 GB Runs comfortably 74K 59.7
Q5_K_M 5.67 8.1 GB 10.2 GB Runs comfortably 82K 68.3
AWQ 4-bit 4.25 7.9 GB 10.0 GB Runs comfortably 83K 69.7
GPTQ 4-bit 4.25 7.9 GB 10.0 GB Runs comfortably 83K 69.7
MXFP4 4.25 7.9 GB 10.0 GB Runs comfortably 83K 69.7
Q5_K_S 5.52 7.9 GB 10.0 GB Runs comfortably 84K 69.9
Q4_K_M 4.85 7.0 GB 9.1 GB Runs comfortably 89K 77.6
Q4_K_S 4.58 6.7 GB 8.8 GB Runs comfortably 91K 81.2
Q4_0 4.55 6.6 GB 8.7 GB Runs comfortably 91K 81.6
IQ4_XS 4.25 6.3 GB 8.3 GB Runs comfortably 94K 86.0
Q3_K_M 3.91 5.8 GB 7.9 GB Runs comfortably 97K 91.7
IQ3_M 3.70 5.6 GB 7.6 GB Runs comfortably 98K 95.5
IQ3_XXS 3.06 4.7 GB 6.8 GB Runs comfortably 104K 110
Q2_K 2.63 4.2 GB 6.3 GB Runs comfortably 107K 122
IQ2_XXS 2.06 3.5 GB 5.6 GB Runs comfortably 112K 142
IQ1_M 1.75 3.1 GB 5.2 GB Runs comfortably 114K 157

Mistral NeMo 12B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 286
A100 80GB 80 2039 Runs comfortably 128K 161
RTX 5090 32 1792 Runs comfortably 128K 149
RTX 5080 16 960 Runs comfortably 41K 81.4
RTX 4090 24 1008 Runs comfortably 89K 77.6
RTX 5070 Ti 16 896 Runs comfortably 41K 76.1
RTX 3090 24 936 Runs comfortably 89K 75.3
Radeon RX 7900 XTX 24 960 Runs comfortably 89K 70.3
L40S 48 864 Runs comfortably 128K 66.7
RTX A6000 48 768 Runs comfortably 128K 62.0
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 61.0
RTX 5070 12 672 Runs comfortably 17K 57.3
RTX 4080 Super 16 736 Runs comfortably 41K 57.0
RTX 4070 Ti Super 16 672 Runs comfortably 41K 52.1
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 45.5
RTX 4070 Super 12 504 Runs comfortably 17K 39.2
RTX 4070 12 504 Runs comfortably 17K 39.2
RTX 5060 Ti 16GB 16 448 Runs comfortably 41K 38.4
RTX 3080 10GB 10 760 Won't fit 5K 34.0
Arc B580 12 456 Runs comfortably 17K 30.1
RTX 3060 12GB 12 360 Runs comfortably 17K 29.3
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 23.5
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 22.9
RTX 4060 Ti 16GB 16 288 Runs comfortably 41K 22.5
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 17.5

Architecture

Parameters12.3B
Layers40
Hidden size5120
Attention heads / KV heads32 / 8
Head dimension128
Vocabulary131,072
Trained context128K
KV cache per 1K tokens0 GB
Hugging Facemistralai/Mistral-Nemo-Instruct-2407

The Mistral family

Dense models — 7B, NeMo 12B, Small 24B and Large 123B — alongside the two Mixtral mixture-of-experts releases and Codestral for code. Vocabulary size is not consistent across the family: 32k on 7B, Large and Codestral, 131k on NeMo and Small, which changes how much of a small model is embedding table.

huggingface.co/mistralai · mistral.ai · all 7 Mistral models

Direct answers