OpenAI on NVIDIA Ada

Can I run gpt-oss 120B-A5.1B on an L40S?

Not at Q4_K_M — it needs 67.4 GB against 44.3 GB available. Drop to IQ3_XXS and it fits, at about 159 tokens per second.

Won't fit

67.4 GB of 44.3 GB · 152%
070 GB
Weights 66.0 GB
KV cache 0.6 GB
Runtime overhead 0.8 GB
Over the limit 23.1 GB

Short by 23.1 GB. You can run it with 23 of 36 layers on the L40S and the rest in system RAM, at roughly 27.8 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation27.8tok/s
Prompt processing7453tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of gpt-oss 120B-A5.1B on a L40S

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 217.6 GB 218.9 GB Won't fit 4.73
INT8 / W8A8 8.50 115.3 GB 116.7 GB Won't fit 10.6
FP8 (E4M3) 8.00 108.8 GB 110.1 GB Won't fit 11.6
Q6_K 6.56 89.2 GB 90.6 GB Won't fit 15.7
Q5_K_M 5.67 77.1 GB 78.5 GB Won't fit 20.6
Q5_K_S 5.52 75.1 GB 76.4 GB Won't fit 21.1
Q4_K_M 4.85 66.0 GB 67.4 GB Won't fit 27.8
Q4_K_S 4.58 62.4 GB 63.8 GB Won't fit 31.1
Q4_0 4.55 62.0 GB 63.4 GB Won't fit 31.2
AWQ 4-bit 4.25 59.4 GB 60.7 GB Won't fit 37.0
GPTQ 4-bit 4.25 59.4 GB 60.7 GB Won't fit 37.0
MXFP4 4.25 59.4 GB 60.7 GB Won't fit 37.0
IQ4_XS 4.25 58.0 GB 59.3 GB Won't fit 37.7
Q3_K_M 3.91 53.4 GB 54.7 GB Won't fit 47.1
IQ3_M 3.70 50.6 GB 51.9 GB Won't fit 58.8
IQ3_XXS 3.06 41.9 GB 43.3 GB Fits, but tight 23K 159
Q2_K 2.63 36.1 GB 37.5 GB Runs comfortably 105K 175
IQ2_XXS 2.06 28.5 GB 29.8 GB Runs comfortably 128K 202
IQ1_M 1.75 24.3 GB 25.7 GB Runs comfortably 128K 220

Also worth checking