Qwen · 235.1B parameters · 22.1B active
Qwen3 235B-A22B VRAM requirements
Qwen3 235B-A22B has 94 layers and uses grouped-query attention (4 KV heads). At Q4_K_M the weights come to 132.8 GB.
Won't fit
135.1 GB of 21.8 GB · 621%0141 GB
Weights
132.8 GB
KV cache
1.5 GB
Runtime overhead
0.8 GB
Over the limit
113.4 GB
Short by 113.4 GB. You can run it with 13 of 94 layers on the RTX 4090 and the rest in system RAM, at roughly 3.23 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.
Generation3.23tok/s
Prompt processing1565tok/s
Max context0tokens
KV per 1K tokens0GB
Every quantisation of Qwen3 235B-A22B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 437.9 GB | 440.2 GB | Won't fit | — | 0.94 |
| INT8 / W8A8 | 8.50 | 232.3 GB | 234.6 GB | Won't fit | — | 1.79 |
| Q8_0 (GGUF) | 8.50 | 232.3 GB | 234.6 GB | Won't fit | — | 1.79 |
| FP8 (E4M3) | 8.00 | 218.9 GB | 221.2 GB | Won't fit | — | 1.92 |
| Q6_K | 6.56 | 179.5 GB | 181.8 GB | Won't fit | — | 2.36 |
| Q5_K_M | 5.67 | 155.2 GB | 157.5 GB | Won't fit | — | 2.74 |
| Q5_K_S | 5.52 | 151.1 GB | 153.4 GB | Won't fit | — | 2.84 |
| Q4_K_M | 4.85 | 132.8 GB | 135.1 GB | Won't fit | — | 3.23 |
| Q4_K_S | 4.58 | 125.5 GB | 127.8 GB | Won't fit | — | 3.44 |
| Q4_0 | 4.55 | 124.7 GB | 127.0 GB | Won't fit | — | 3.46 |
| AWQ 4-bit | 4.25 | 118.0 GB | 120.3 GB | Won't fit | — | 3.68 |
| GPTQ 4-bit | 4.25 | 118.0 GB | 120.3 GB | Won't fit | — | 3.68 |
| MXFP4 | 4.25 | 118.0 GB | 120.3 GB | Won't fit | — | 3.68 |
| IQ4_XS | 4.25 | 116.5 GB | 118.8 GB | Won't fit | — | 3.72 |
| Q3_K_M | 3.91 | 107.2 GB | 109.5 GB | Won't fit | — | 4.10 |
| IQ3_M | 3.70 | 101.5 GB | 103.8 GB | Won't fit | — | 4.36 |
| IQ3_XXS | 3.06 | 84.1 GB | 86.4 GB | Won't fit | — | 5.34 |
| Q2_K | 2.63 | 72.4 GB | 74.7 GB | Won't fit | — | 6.38 |
| IQ2_XXS | 2.06 | 56.9 GB | 59.2 GB | Won't fit | — | 8.55 |
| IQ1_M | 1.75 | 48.4 GB | 50.7 GB | Won't fit | — | 10.4 |
Qwen3 235B-A22B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 128K | 24.7 |
| Mac Studio M4 Max 128GB | 128 | 546 | Won't fit | — | 7.42 |
| H100 SXM 80GB | 80 | 3350 | Won't fit | — | 6.80 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Won't fit | — | 6.20 |
| A100 80GB | 80 | 2039 | Won't fit | — | 6.02 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Won't fit | — | 4.85 |
| RTX A6000 | 48 | 768 | Won't fit | — | 4.05 |
| L40S | 48 | 864 | Won't fit | — | 3.90 |
| RTX 5090 | 32 | 1792 | Won't fit | — | 3.83 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Won't fit | — | 3.70 |
| RTX 5080 | 16 | 960 | Won't fit | — | 3.36 |
| RTX 3090 | 24 | 936 | Won't fit | — | 3.36 |
| RTX 5070 Ti | 16 | 896 | Won't fit | — | 3.36 |
| RTX 5060 Ti 16GB | 16 | 448 | Won't fit | — | 3.34 |
| RTX 5070 | 12 | 672 | Won't fit | — | 3.26 |
| RTX 4090 | 24 | 1008 | Won't fit | — | 3.23 |
| RTX 3060 12GB | 12 | 360 | Won't fit | — | 3.07 |
| Radeon RX 7900 XTX | 24 | 960 | Won't fit | — | 3.07 |
| RTX 3080 10GB | 10 | 760 | Won't fit | — | 3.06 |
| RTX 4080 Super | 16 | 736 | Won't fit | — | 3.05 |
| RTX 4070 Ti Super | 16 | 672 | Won't fit | — | 3.05 |
| RTX 4060 Ti 16GB | 16 | 288 | Won't fit | — | 3.01 |
| RTX 4070 Super | 12 | 504 | Won't fit | — | 2.96 |
| RTX 4070 | 12 | 504 | Won't fit | — | 2.96 |
| Arc B580 | 12 | 456 | Won't fit | — | 2.51 |
Architecture
| Parameters | 235.1B |
| Active per token | 22.1B of 128 experts, top-8 |
| Layers | 94 |
| Hidden size | 4096 |
| Attention heads / KV heads | 64 / 4 |
| Head dimension | 128 |
| Vocabulary | 151,936 |
| Trained context | 128K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | Qwen/Qwen3-235B-A22B |
Direct answers
Qwen3 235B-A22B on RTX 5090
See the verdictQwen3 235B-A22B on RTX 4090
See the verdictQwen3 235B-A22B on RTX 3090
See the verdictQwen3 235B-A22B on RTX 5080
See the verdictQwen3 235B-A22B on RTX 5070 Ti
See the verdictQwen3 235B-A22B on RTX 5070
See the verdictQwen3 235B-A22B on RTX 4070 Ti Super
See the verdictQwen3 235B-A22B on RTX 4070 Super
See the verdictQwen3 235B-A22B on RTX 4070
See the verdict
See the verdictQwen3 235B-A22B on RTX 4090
See the verdictQwen3 235B-A22B on RTX 3090
See the verdictQwen3 235B-A22B on RTX 5080
See the verdictQwen3 235B-A22B on RTX 5070 Ti
See the verdictQwen3 235B-A22B on RTX 5070
See the verdictQwen3 235B-A22B on RTX 4070 Ti Super
See the verdictQwen3 235B-A22B on RTX 4070 Super
See the verdictQwen3 235B-A22B on RTX 4070
See the verdict