NVIDIA · Ada

What runs on a RTX 4060

8 GB at 272 GB/s. The largest model that fits at Q4_K_M is GLM-4 9B, at roughly 29.1 tokens per second.

Fits, but tight

6.4 GB of 6.7 GB · 96%
07 GB
Weights 4.6 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

This fits with almost nothing to spare. A background application claiming VRAM will push it over. Drop to the next quantisation down, or quantise the KV cache to Q8_0 — that halves the cache for no meaningful quality loss.

Generation31.7tok/s
Prompt processing785tok/s
Max context10Ktokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 8B on a RTX 4060

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 15.0 GB 16.8 GB Won't fit 3.25
INT8 / W8A8 8.50 7.7 GB 9.5 GB Won't fit 8.92
Q8_0 (GGUF) 8.50 7.7 GB 9.5 GB Won't fit 8.92
FP8 (E4M3) 8.00 7.5 GB 9.3 GB Won't fit 9.18
Q6_K 6.56 6.1 GB 8.0 GB Won't fit 14.3
AWQ 4-bit 4.25 5.4 GB 7.2 GB Won't fit 4K 19.5
GPTQ 4-bit 4.25 5.4 GB 7.2 GB Won't fit 4K 19.5
MXFP4 4.25 5.4 GB 7.2 GB Won't fit 4K 19.5
Q5_K_M 5.67 5.3 GB 7.1 GB Won't fit 5K 21.4
Q5_K_S 5.52 5.2 GB 7.0 GB Won't fit 6K 23.8
Q4_K_M 4.85 4.6 GB 6.4 GB Fits, but tight 10K 31.7
Q4_K_S 4.58 4.4 GB 6.2 GB Fits, but tight 12K 33.1
Q4_0 4.55 4.4 GB 6.2 GB Fits, but tight 12K 33.3
IQ4_XS 4.25 4.1 GB 5.9 GB Runs comfortably 14K 35.0
Q3_K_M 3.91 3.8 GB 5.7 GB Runs comfortably 16K 37.2
IQ3_M 3.70 3.7 GB 5.5 GB Runs comfortably 18K 38.7
IQ3_XXS 3.06 3.2 GB 5.0 GB Runs comfortably 22K 44.2
Q2_K 2.63 2.8 GB 4.6 GB Runs comfortably 25K 48.8
IQ2_XXS 2.06 2.3 GB 4.2 GB Runs comfortably 28K 56.7
IQ1_M 1.75 2.1 GB 3.9 GB Runs comfortably 30K 62.1

Models on a RTX 4060

Q4_K_M at 8K context. Bandwidth sets the speed; capacity sets the ceiling.

ModelParamsWeightsVerdictMax ctxtok/s
DeepSeek-R1 671B-A37B 671B 379.0 GB Won't fit 1.79
Qwen3 235B-A22B 235.1B 132.8 GB Won't fit 2.89
gpt-oss 120B-A5.1B 116.8B 66.0 GB Won't fit 12.4
Llama 4 Scout 109B-A17B 109B 61.7 GB Won't fit 3.86
Qwen2.5 72B 72.7B 41.2 GB Won't fit 0.96
Llama 3.1 70B 70.5B 40.0 GB Won't fit 0.98
Mixtral 8x7B 46.7B 26.4 GB Won't fit 5.52
Command R 35B 35.0B 20.1 GB Won't fit 1.53
Yi-1.5 34B 34.4B 19.5 GB Won't fit 2.21
Qwen3 32B 32.8B 18.6 GB Won't fit 2.31
Qwen2.5-Coder 32B 32.8B 18.6 GB Won't fit 2.31
DeepSeek-R1-Distill-Qwen 32B 32.8B 18.6 GB Won't fit 2.31
Qwen3 30B-A3B 30.5B 17.3 GB Won't fit 19.7
Gemma 3 27B 27.4B 15.6 GB Won't fit 3.06
Mistral Small 24B 23.6B 13.4 GB Won't fit 3.63
gpt-oss 20B-A3.6B 20.9B 11.9 GB Won't fit 25.0
Qwen3 14B 14.8B 8.5 GB Won't fit 7.03
Phi-4 14B 14.7B 8.4 GB Won't fit 6.76
Mistral NeMo 12B 12.3B 7.0 GB Won't fit 9.94
Gemma 3 12B 12.2B 7.0 GB Won't fit 11.5
GLM-4 9B 9.4B 5.4 GB Fits, but tight 13K 29.1
Qwen3 8B 8.2B 4.7 GB Fits, but tight 8K 30.7
Llama 3.1 8B 8.0B 4.6 GB Fits, but tight 10K 31.7
DeepSeek-R1-Distill-Llama 8B 8.0B 4.6 GB Fits, but tight 10K 31.7
Qwen2.5 7B 7.6B 4.4 GB Runs comfortably 28K 35.2
Mistral 7B v0.3 7.3B 4.1 GB Runs comfortably 14K 35.1
Gemma 3 4B 4.3B 2.5 GB Runs comfortably 128K 60.1
Qwen3 4B 4.0B 2.3 GB Runs comfortably 26K 56.3
Llama 3.2 3B 3.2B 1.8 GB Runs comfortably 37K 70.7
Llama 3.2 1B 1.2B 0.7 GB Runs comfortably 128K 190

Direct answers