OpenAI · gpt-oss · 20.9B parameters · 3.6B active

gpt-oss 20B-A3.6B VRAM requirements

gpt-oss 20B-A3.6B has 24 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 11.9 GB, and the best quantisation that fits a 24 GB card is INT8 / W8A8.

Runs comfortably

13.1 GB of 21.8 GB · 60%
022 GB
Weights 11.9 GB
KV cache 0.4 GB
Runtime overhead 0.8 GB

gpt-oss 20B-A3.6B at Q4_K_M leaves 8.7 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation188tok/s
Prompt processing9625tok/s
Max context128Ktokens
KV per 1K tokens0GB

Every quantisation of gpt-oss 20B-A3.6B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 38.9 GB 40.1 GB Won't fit 10.2
INT8 / W8A8 8.50 20.4 GB 21.6 GB Fits, but tight 12K 123
FP8 (E4M3) 8.00 19.5 GB 20.6 GB Fits, but tight 32K 128
Q6_K 6.56 16.0 GB 17.1 GB Runs comfortably 107K 150
Q5_K_M 5.67 13.8 GB 15.0 GB Runs comfortably 128K 168
Q5_K_S 5.52 13.4 GB 14.6 GB Runs comfortably 128K 172
AWQ 4-bit 4.25 11.9 GB 13.1 GB Runs comfortably 128K 188
GPTQ 4-bit 4.25 11.9 GB 13.1 GB Runs comfortably 128K 188
MXFP4 4.25 11.9 GB 13.1 GB Runs comfortably 128K 188
Q4_K_M 4.85 11.9 GB 13.1 GB Runs comfortably 128K 188
Q4_K_S 4.58 11.3 GB 12.4 GB Runs comfortably 128K 196
Q4_0 4.55 11.2 GB 12.4 GB Runs comfortably 128K 196
IQ4_XS 4.25 10.5 GB 11.7 GB Runs comfortably 128K 206
Q3_K_M 3.91 9.7 GB 10.9 GB Runs comfortably 128K 217
IQ3_M 3.70 9.2 GB 10.4 GB Runs comfortably 128K 225
IQ3_XXS 3.06 7.8 GB 8.9 GB Runs comfortably 128K 253
Q2_K 2.63 6.8 GB 8.0 GB Runs comfortably 128K 276
IQ2_XXS 2.06 5.5 GB 6.7 GB Runs comfortably 128K 313
IQ1_M 1.75 4.8 GB 5.9 GB Runs comfortably 128K 338

gpt-oss 20B-A3.6B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 470
A100 80GB 80 2039 Runs comfortably 128K 327
RTX 5090 32 1792 Runs comfortably 128K 311
RTX 5080 16 960 Fits, but tight 33K 195
RTX 4090 24 1008 Runs comfortably 128K 188
RTX 5070 Ti 16 896 Fits, but tight 33K 185
RTX 3090 24 936 Runs comfortably 128K 183
Radeon RX 7900 XTX 24 960 Runs comfortably 128K 173
L40S 48 864 Runs comfortably 128K 166
RTX A6000 48 768 Runs comfortably 128K 156
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 154
RTX 4080 Super 16 736 Fits, but tight 33K 145
RTX 4070 Ti Super 16 672 Fits, but tight 33K 134
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 119
RTX 5060 Ti 16GB 16 448 Fits, but tight 33K 102
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 64.8
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 63.2
RTX 4060 Ti 16GB 16 288 Fits, but tight 33K 62.2
RTX 5070 12 672 Won't fit 53.6
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 49.1
RTX 4070 Super 12 504 Won't fit 45.3
RTX 4070 12 504 Won't fit 45.3
RTX 3060 12GB 12 360 Won't fit 42.3
Arc B580 12 456 Won't fit 37.5
RTX 3080 10GB 10 760 Won't fit 36.3

Architecture

Parameters20.9B
Active per token3.6B of 32 experts, top-4
Layers24
Hidden size2880
Attention heads / KV heads64 / 8
Head dimension64
Vocabulary201,088
Trained context128K
KV cache per 1K tokens0 GB
Hugging Faceopenai/gpt-oss-20b

The OpenAI family

gpt-oss 20B and 120B are mixture-of-experts models that ship natively in MXFP4, so the quantisation ladder starts at the published 4-bit weights rather than at fp16. Both read very few parameters per token — 3.6B and 5.1B — which makes them unusually fast for their size.

huggingface.co/openai · github.com/openai/gpt-oss · all 2 OpenAI models

Direct answers