Qwen · 30.5B parameters · 3.3B active

Qwen3 30B-A3B VRAM requirements

Qwen3 30B-A3B has 48 layers and uses grouped-query attention (4 KV heads). At Q4_K_M the weights come to 17.3 GB, and the best quantisation that fits a 24 GB card is Q5_K_M.

Runs comfortably

18.8 GB of 21.8 GB · 86%
022 GB
Weights 17.3 GB
KV cache 0.8 GB
Runtime overhead 0.8 GB

Qwen3 30B-A3B at Q4_K_M leaves 2.9 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation118tok/s
Prompt processing10374tok/s
Max context39Ktokens
KV per 1K tokens0GB

Every quantisation of Qwen3 30B-A3B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 56.9 GB 58.4 GB Won't fit 8.34
INT8 / W8A8 8.50 30.1 GB 31.6 GB Won't fit 24.6
Q8_0 (GGUF) 8.50 30.1 GB 31.6 GB Won't fit 24.6
FP8 (E4M3) 8.00 28.4 GB 30.0 GB Won't fit 28.3
Q6_K 6.56 23.3 GB 24.9 GB Won't fit 49.9
Q5_K_M 5.67 20.2 GB 21.7 GB Fits, but tight 9K 111
Q5_K_S 5.52 19.6 GB 21.2 GB Fits, but tight 14K 112
Q4_K_M 4.85 17.3 GB 18.8 GB Runs comfortably 39K 118
Q4_K_S 4.58 16.3 GB 17.9 GB Runs comfortably 49K 121
Q4_0 4.55 16.2 GB 17.8 GB Runs comfortably 50K 122
AWQ 4-bit 4.25 16.0 GB 17.5 GB Runs comfortably 54K 123
GPTQ 4-bit 4.25 16.0 GB 17.5 GB Runs comfortably 54K 123
MXFP4 4.25 16.0 GB 17.5 GB Runs comfortably 54K 123
IQ4_XS 4.25 15.2 GB 16.7 GB Runs comfortably 62K 125
Q3_K_M 3.91 14.0 GB 15.5 GB Runs comfortably 74K 129
IQ3_M 3.70 13.3 GB 14.8 GB Runs comfortably 82K 132
IQ3_XXS 3.06 11.1 GB 12.6 GB Runs comfortably 106K 140
Q2_K 2.63 9.6 GB 11.1 GB Runs comfortably 122K 147
IQ2_XXS 2.06 7.6 GB 9.1 GB Runs comfortably 128K 156
IQ1_M 1.75 6.5 GB 8.0 GB Runs comfortably 128K 162

Qwen3 30B-A3B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 192
A100 80GB 80 2039 Runs comfortably 128K 163
RTX 5090 32 1792 Runs comfortably 120K 159
RTX 4090 24 1008 Runs comfortably 39K 118
RTX 3090 24 936 Runs comfortably 39K 117
Radeon RX 7900 XTX 24 960 Runs comfortably 39K 112
L40S 48 864 Runs comfortably 128K 109
RTX A6000 48 768 Runs comfortably 128K 105
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 104
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 86.4
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 53.5
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 52.4
RTX 5080 16 960 Won't fit 46.2
RTX 5070 Ti 16 896 Won't fit 45.7
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 42.2
RTX 4080 Super 16 736 Won't fit 40.9
RTX 4070 Ti Super 16 672 Won't fit 40.2
RTX 5060 Ti 16GB 16 448 Won't fit 39.8
RTX 4060 Ti 16GB 16 288 Won't fit 32.0
RTX 5070 12 672 Won't fit 29.5
RTX 4070 Super 12 504 Won't fit 26.3
RTX 4070 12 504 Won't fit 26.3
RTX 3060 12GB 12 360 Won't fit 26.1
RTX 3080 10GB 10 760 Won't fit 24.7
Arc B580 12 456 Won't fit 22.4

Architecture

Parameters30.5B
Active per token3.3B of 128 experts, top-8
Layers48
Hidden size2048
Attention heads / KV heads32 / 4
Head dimension128
Vocabulary151,936
Trained context128K
KV cache per 1K tokens0 GB
Hugging FaceQwen/Qwen3-30B-A3B

Direct answers