Microsoft · Phi-4 · 14.7B parameters

Phi-4 14B VRAM requirements

Phi-4 14B has 40 layers and uses grouped-query attention (10 KV heads). At Q4_K_M the weights come to 8.4 GB, and the best quantisation that fits a 24 GB card is INT8 / W8A8.

Runs comfortably

10.8 GB of 21.8 GB · 50%
022 GB
Weights 8.4 GB
KV cache 1.6 GB
Runtime overhead 0.8 GB

Phi-4 14B at Q4_K_M leaves 11.0 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation65.0tok/s
Prompt processing2357tok/s
Max context16Ktokens
KV per 1K tokens0GB

Every quantisation of Phi-4 14B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 27.4 GB 29.8 GB Won't fit 3.94
INT8 / W8A8 8.50 14.3 GB 16.7 GB Runs comfortably 16K 39.7
Q8_0 (GGUF) 8.50 14.3 GB 16.7 GB Runs comfortably 16K 39.7
FP8 (E4M3) 8.00 13.7 GB 16.1 GB Runs comfortably 16K 41.4
Q6_K 6.56 11.2 GB 13.6 GB Runs comfortably 16K 49.7
Q5_K_M 5.67 9.7 GB 12.1 GB Runs comfortably 16K 56.9
Q5_K_S 5.52 9.4 GB 11.8 GB Runs comfortably 16K 58.3
AWQ 4-bit 4.25 8.7 GB 11.1 GB Runs comfortably 16K 62.9
GPTQ 4-bit 4.25 8.7 GB 11.1 GB Runs comfortably 16K 62.9
MXFP4 4.25 8.7 GB 11.1 GB Runs comfortably 16K 62.9
Q4_K_M 4.85 8.4 GB 10.8 GB Runs comfortably 16K 65.0
Q4_K_S 4.58 7.9 GB 10.3 GB Runs comfortably 16K 68.1
Q4_0 4.55 7.9 GB 10.3 GB Runs comfortably 16K 68.5
IQ4_XS 4.25 7.4 GB 9.8 GB Runs comfortably 16K 72.4
Q3_K_M 3.91 6.9 GB 9.3 GB Runs comfortably 16K 77.4
IQ3_M 3.70 6.5 GB 8.9 GB Runs comfortably 16K 80.9
IQ3_XXS 3.06 5.5 GB 7.9 GB Runs comfortably 16K 93.6
Q2_K 2.63 4.8 GB 7.2 GB Runs comfortably 16K 105
IQ2_XXS 2.06 3.9 GB 6.3 GB Runs comfortably 16K 124
IQ1_M 1.75 3.4 GB 5.8 GB Runs comfortably 16K 138

Phi-4 14B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 16K 241
A100 80GB 80 2039 Runs comfortably 16K 135
RTX 5090 32 1792 Runs comfortably 16K 125
RTX 5080 16 960 Runs comfortably 16K 68.1
RTX 4090 24 1008 Runs comfortably 16K 65.0
RTX 5070 Ti 16 896 Runs comfortably 16K 63.7
RTX 3090 24 936 Runs comfortably 16K 63.0
Radeon RX 7900 XTX 24 960 Runs comfortably 16K 58.8
L40S 48 864 Runs comfortably 16K 55.8
RTX A6000 48 768 Runs comfortably 16K 51.8
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 16K 51.0
RTX 4080 Super 16 736 Runs comfortably 16K 47.6
RTX 4070 Ti Super 16 672 Runs comfortably 16K 43.5
Mac Studio M4 Max 128GB 128 546 Runs comfortably 16K 38.0
RTX 5070 12 672 Won't fit 6K 32.6
RTX 5060 Ti 16GB 16 448 Runs comfortably 16K 32.1
RTX 4070 Super 12 504 Won't fit 6K 24.4
RTX 4070 12 504 Won't fit 6K 24.4
RTX 3060 12GB 12 360 Won't fit 6K 19.9
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 16K 19.6
Arc B580 12 456 Won't fit 6K 19.2
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 16K 19.1
RTX 4060 Ti 16GB 16 288 Runs comfortably 16K 18.8
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 16K 14.6
RTX 3080 10GB 10 760 Won't fit 12.9

Architecture

Parameters14.7B
Layers40
Hidden size5120
Attention heads / KV heads40 / 10
Head dimension128
Vocabulary100,352
Trained context16K
KV cache per 1K tokens0 GB
Hugging Facemicrosoft/phi-4

The Microsoft family

The Phi models are trained on curated and synthetic data to punch above their parameter count. Phi-4 14B keeps a 16K context — short by current standards, but easy on the cache. Phi-4-mini goes to 128K with a 200k vocabulary.

huggingface.co/microsoft · all 2 Microsoft models

Direct answers