OpenAI · gpt-oss · 116.8B parameters · 5.1B active

gpt-oss 120B-A5.1B VRAM requirements

gpt-oss 120B-A5.1B has 36 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 66.0 GB.

Won't fit

67.4 GB of 21.8 GB · 310%
070 GB
Weights 66.0 GB
KV cache 0.6 GB
Runtime overhead 0.8 GB
Over the limit 45.6 GB

Short by 45.6 GB. You can run it with 11 of 36 layers on the RTX 4090 and the rest in system RAM, at roughly 16.4 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation16.4tok/s
Prompt processing6794tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of gpt-oss 120B-A5.1B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 217.6 GB 218.9 GB Won't fit 4.21
INT8 / W8A8 8.50 115.3 GB 116.7 GB Won't fit 8.40
FP8 (E4M3) 8.00 108.8 GB 110.1 GB Won't fit 8.88
Q6_K 6.56 89.2 GB 90.6 GB Won't fit 11.3
Q5_K_M 5.67 77.1 GB 78.5 GB Won't fit 13.4
Q5_K_S 5.52 75.1 GB 76.4 GB Won't fit 13.7
Q4_K_M 4.85 66.0 GB 67.4 GB Won't fit 16.4
Q4_K_S 4.58 62.4 GB 63.8 GB Won't fit 17.3
Q4_0 4.55 62.0 GB 63.4 GB Won't fit 17.4
AWQ 4-bit 4.25 59.4 GB 60.7 GB Won't fit 18.7
GPTQ 4-bit 4.25 59.4 GB 60.7 GB Won't fit 18.7
MXFP4 4.25 59.4 GB 60.7 GB Won't fit 18.7
IQ4_XS 4.25 58.0 GB 59.3 GB Won't fit 19.1
Q3_K_M 3.91 53.4 GB 54.7 GB Won't fit 21.3
IQ3_M 3.70 50.6 GB 51.9 GB Won't fit 23.2
IQ3_XXS 3.06 41.9 GB 43.3 GB Won't fit 30.7
Q2_K 2.63 36.1 GB 37.5 GB Won't fit 39.8
IQ2_XXS 2.06 28.5 GB 29.8 GB Won't fit 63.5
IQ1_M 1.75 24.3 GB 25.7 GB Won't fit 105

gpt-oss 120B-A5.1B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Fits, but tight 108K 323
A100 80GB 80 2039 Fits, but tight 108K 227
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 107
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 83.3
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 45.6
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 34.6
RTX A6000 48 768 Won't fit 28.6
L40S 48 864 Won't fit 27.8
RTX 5090 32 1792 Won't fit 21.5
Mac Mini M4 Pro 48GB 48 273 Won't fit 20.4
RTX 3090 24 936 Won't fit 17.1
RTX 4090 24 1008 Won't fit 16.4
RTX 5080 16 960 Won't fit 15.9
RTX 5070 Ti 16 896 Won't fit 15.8
Radeon RX 7900 XTX 24 960 Won't fit 15.6
RTX 5060 Ti 16GB 16 448 Won't fit 15.5
RTX 5070 12 672 Won't fit 14.5
RTX 4080 Super 16 736 Won't fit 14.3
RTX 4070 Ti Super 16 672 Won't fit 14.3
RTX 4060 Ti 16GB 16 288 Won't fit 13.8
RTX 3060 12GB 12 360 Won't fit 13.6
RTX 3080 10GB 10 760 Won't fit 13.4
RTX 4070 Super 12 504 Won't fit 13.1
RTX 4070 12 504 Won't fit 13.1
Arc B580 12 456 Won't fit 11.1

Architecture

Parameters116.8B
Active per token5.1B of 128 experts, top-4
Layers36
Hidden size2880
Attention heads / KV heads64 / 8
Head dimension64
Vocabulary201,088
Trained context128K
KV cache per 1K tokens0 GB
Hugging Faceopenai/gpt-oss-120b

The OpenAI family

gpt-oss 20B and 120B are mixture-of-experts models that ship natively in MXFP4, so the quantisation ladder starts at the published 4-bit weights rather than at fp16. Both read very few parameters per token — 3.6B and 5.1B — which makes them unusually fast for their size.

huggingface.co/openai · github.com/openai/gpt-oss · all 2 OpenAI models

Direct answers