Mistral · Codestral · 22.2B parameters

Codestral 22B VRAM requirements

Codestral 22B has 56 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 12.6 GB, and the best quantisation that fits a 24 GB card is Q6_K.

Runs comfortably

15.2 GB of 21.8 GB · 70%
022 GB
Weights 12.6 GB
KV cache 1.8 GB
Runtime overhead 0.9 GB

Codestral 22B at Q4_K_M leaves 6.6 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation44.3tok/s
Prompt processing1561tok/s
Max context32Ktokens
KV per 1K tokens0GB

Every quantisation of Codestral 22B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 41.4 GB 44.0 GB Won't fit 1.56
INT8 / W8A8 8.50 21.9 GB 24.5 GB Won't fit 9.30
Q8_0 (GGUF) 8.50 21.9 GB 24.5 GB Won't fit 9.30
FP8 (E4M3) 8.00 20.7 GB 23.3 GB Won't fit 1K 12.0
Q6_K 6.56 17.0 GB 19.6 GB Runs comfortably 18K 33.5
Q5_K_M 5.67 14.7 GB 17.3 GB Runs comfortably 29K 38.4
Q5_K_S 5.52 14.3 GB 16.9 GB Runs comfortably 30K 39.4
Q4_K_M 4.85 12.6 GB 15.2 GB Runs comfortably 32K 44.3
Q4_K_S 4.58 11.9 GB 14.5 GB Runs comfortably 32K 46.6
Q4_0 4.55 11.8 GB 14.4 GB Runs comfortably 32K 46.9
AWQ 4-bit 4.25 11.5 GB 14.1 GB Runs comfortably 32K 47.9
GPTQ 4-bit 4.25 11.5 GB 14.1 GB Runs comfortably 32K 47.9
MXFP4 4.25 11.5 GB 14.1 GB Runs comfortably 32K 47.9
IQ4_XS 4.25 11.0 GB 13.6 GB Runs comfortably 32K 49.9
Q3_K_M 3.91 10.2 GB 12.8 GB Runs comfortably 32K 53.7
IQ3_M 3.70 9.6 GB 12.3 GB Runs comfortably 32K 56.4
IQ3_XXS 3.06 8.0 GB 10.6 GB Runs comfortably 32K 66.4
Q2_K 2.63 6.9 GB 9.5 GB Runs comfortably 32K 75.5
IQ2_XXS 2.06 5.5 GB 8.1 GB Runs comfortably 32K 92.1
IQ1_M 1.75 4.7 GB 7.3 GB Runs comfortably 32K 105

Codestral 22B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 32K 165
A100 80GB 80 2039 Runs comfortably 32K 92.0
RTX 5090 32 1792 Runs comfortably 32K 85.6
RTX 4090 24 1008 Runs comfortably 32K 44.3
RTX 3090 24 936 Runs comfortably 32K 43.0
Radeon RX 7900 XTX 24 960 Runs comfortably 32K 40.1
L40S 48 864 Runs comfortably 32K 38.1
RTX A6000 48 768 Runs comfortably 32K 35.3
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 32K 34.8
Mac Studio M4 Max 128GB 128 546 Runs comfortably 32K 25.9
RTX 5080 16 960 Won't fit 4K 20.8
RTX 5070 Ti 16 896 Won't fit 4K 20.2
RTX 4080 Super 16 736 Won't fit 4K 16.8
RTX 4070 Ti Super 16 672 Won't fit 4K 16.1
RTX 5060 Ti 16GB 16 448 Won't fit 4K 14.3
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 32K 13.4
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 32K 13.0
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 32K 9.97
RTX 4060 Ti 16GB 16 288 Won't fit 4K 9.75
RTX 5070 12 672 Won't fit 7.21
RTX 4070 Super 12 504 Won't fit 6.26
RTX 4070 12 504 Won't fit 6.26
RTX 3060 12GB 12 360 Won't fit 6.11
Arc B580 12 456 Won't fit 5.21
RTX 3080 10GB 10 760 Won't fit 5.16

Architecture

Parameters22.2B
Layers56
Hidden size6144
Attention heads / KV heads48 / 8
Head dimension128
Vocabulary32,768
Trained context32K
KV cache per 1K tokens0 GB
Hugging Facemistralai/Codestral-22B-v0.1

The Mistral family

Dense models — 7B, NeMo 12B, Small 24B and Large 123B — alongside the two Mixtral mixture-of-experts releases and Codestral for code. Vocabulary size is not consistent across the family: 32k on 7B, Large and Codestral, 131k on NeMo and Small, which changes how much of a small model is embedding table.

huggingface.co/mistralai · mistral.ai · all 7 Mistral models

Direct answers