Qwen on NVIDIA Blackwell

Can I run Qwen3 235B-A22B on an RTX 5090?

Not at Q4_K_M — it needs 135.1 GB against 29.3 GB available. You would need 5 of these cards.

Won't fit

135.1 GB of 29.3 GB · 461%
0141 GB
Weights 132.8 GB
KV cache 1.5 GB
Runtime overhead 0.8 GB
Over the limit 105.8 GB

Short by 105.8 GB. You can run it with 19 of 94 layers on the RTX 5090 and the rest in system RAM, at roughly 3.83 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation3.83tok/s
Prompt processing1982tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Qwen3 235B-A22B on a RTX 5090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 437.9 GB 440.2 GB Won't fit 1.05
INT8 / W8A8 8.50 232.3 GB 234.6 GB Won't fit 2.04
Q8_0 (GGUF) 8.50 232.3 GB 234.6 GB Won't fit 2.04
FP8 (E4M3) 8.00 218.9 GB 221.2 GB Won't fit 2.19
Q6_K 6.56 179.5 GB 181.8 GB Won't fit 2.73
Q5_K_M 5.67 155.2 GB 157.5 GB Won't fit 3.20
Q5_K_S 5.52 151.1 GB 153.4 GB Won't fit 3.28
Q4_K_M 4.85 132.8 GB 135.1 GB Won't fit 3.83
Q4_K_S 4.58 125.5 GB 127.8 GB Won't fit 4.08
Q4_0 4.55 124.7 GB 127.0 GB Won't fit 4.11
AWQ 4-bit 4.25 118.0 GB 120.3 GB Won't fit 4.37
GPTQ 4-bit 4.25 118.0 GB 120.3 GB Won't fit 4.37
MXFP4 4.25 118.0 GB 120.3 GB Won't fit 4.37
IQ4_XS 4.25 116.5 GB 118.8 GB Won't fit 4.42
Q3_K_M 3.91 107.2 GB 109.5 GB Won't fit 4.89
IQ3_M 3.70 101.5 GB 103.8 GB Won't fit 5.20
IQ3_XXS 3.06 84.1 GB 86.4 GB Won't fit 6.65
Q2_K 2.63 72.4 GB 74.7 GB Won't fit 8.12
IQ2_XXS 2.06 56.9 GB 59.2 GB Won't fit 11.4
IQ1_M 1.75 48.4 GB 50.7 GB Won't fit 14.9

Also worth checking