HuggingFace · SmolLM2 · 1.7B parameters

SmolLM2 1.7B VRAM requirements

SmolLM2 1.7B has 24 layers and uses full multi-head attention — no GQA, so the cache is large. At Q4_K_M the weights come to 1.0 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

3.3 GB of 21.8 GB · 15%
022 GB
Weights 1.0 GB
KV cache 1.5 GB
Runtime overhead 0.8 GB

SmolLM2 1.7B at Q4_K_M leaves 18.5 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation334tok/s
Prompt processing20263tok/s
Max context8Ktokens
KV per 1K tokens0GB

Every quantisation of SmolLM2 1.7B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 3.2 GB 5.5 GB Runs comfortably 8K 150
INT8 / W8A8 8.50 1.7 GB 4.0 GB Runs comfortably 8K 241
Q8_0 (GGUF) 8.50 1.7 GB 4.0 GB Runs comfortably 8K 241
FP8 (E4M3) 8.00 1.6 GB 3.9 GB Runs comfortably 8K 249
Q6_K 6.56 1.3 GB 3.6 GB Runs comfortably 8K 282
Q5_K_M 5.67 1.1 GB 3.4 GB Runs comfortably 8K 308
Q5_K_S 5.52 1.1 GB 3.4 GB Runs comfortably 8K 312
AWQ 4-bit 4.25 1.0 GB 3.3 GB Runs comfortably 8K 332
GPTQ 4-bit 4.25 1.0 GB 3.3 GB Runs comfortably 8K 332
MXFP4 4.25 1.0 GB 3.3 GB Runs comfortably 8K 332
Q4_K_M 4.85 1.0 GB 3.3 GB Runs comfortably 8K 334
Q4_K_S 4.58 0.9 GB 3.2 GB Runs comfortably 8K 344
Q4_0 4.55 0.9 GB 3.2 GB Runs comfortably 8K 345
IQ4_XS 4.25 0.9 GB 3.1 GB Runs comfortably 8K 356
Q3_K_M 3.91 0.8 GB 3.1 GB Runs comfortably 8K 370
IQ3_M 3.70 0.8 GB 3.0 GB Runs comfortably 8K 379
IQ3_XXS 3.06 0.6 GB 2.9 GB Runs comfortably 8K 410
Q2_K 2.63 0.6 GB 2.8 GB Runs comfortably 8K 434
IQ2_XXS 2.06 0.5 GB 2.7 GB Runs comfortably 8K 470
IQ1_M 1.75 0.4 GB 2.7 GB Runs comfortably 8K 492

SmolLM2 1.7B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 8K 1138
A100 80GB 80 2039 Runs comfortably 8K 669
RTX 5090 32 1792 Runs comfortably 8K 625
RTX 5080 16 960 Runs comfortably 8K 350
RTX 4090 24 1008 Runs comfortably 8K 334
RTX 5070 Ti 16 896 Runs comfortably 8K 327
RTX 3090 24 936 Runs comfortably 8K 324
Radeon RX 7900 XTX 24 960 Runs comfortably 8K 303
L40S 48 864 Runs comfortably 8K 288
RTX A6000 48 768 Runs comfortably 8K 268
RTX 3080 10GB 10 760 Runs comfortably 8K 266
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 8K 264
RTX 5070 12 672 Runs comfortably 8K 249
RTX 4080 Super 16 736 Runs comfortably 8K 247
RTX 4070 Ti Super 16 672 Runs comfortably 8K 226
Mac Studio M4 Max 128GB 128 546 Runs comfortably 8K 198
RTX 4070 Super 12 504 Runs comfortably 8K 171
RTX 4070 12 504 Runs comfortably 8K 171
RTX 5060 Ti 16GB 16 448 Runs comfortably 8K 168
Arc B580 12 456 Runs comfortably 8K 132
RTX 3060 12GB 12 360 Runs comfortably 8K 128
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 8K 103
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 8K 100
RTX 4060 Ti 16GB 16 288 Runs comfortably 8K 98.8
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 8K 77.2

Architecture

Parameters1.7B
Layers24
Hidden size2048
Attention heads / KV heads32 / 32
Head dimension64
Vocabulary49,152
Trained context8K
KV cache per 1K tokens0 GB
Hugging FaceHuggingFaceTB/SmolLM2-1.7B-Instruct

The HuggingFace family

SmolLM2 comes from Hugging Face's own team and targets the small end — 1.7B parameters, a 49k vocabulary and full multi-head attention. It runs on nearly anything.

huggingface.co/HuggingFaceTB · github.com/huggingface/smollm

Direct answers