NVIDIA · Ampere

What runs on a RTX 3070

8 GB at 448 GB/s. The largest model that fits at Q4_K_M is GLM-4 9B, at roughly 49.9 tokens per second.

Fits, but tight

6.4 GB of 6.7 GB · 96%
07 GB
Weights 4.6 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

This fits with almost nothing to spare. A background application claiming VRAM will push it over. Drop to the next quantisation down, or quantise the KV cache to Q8_0 — that halves the cache for no meaningful quality loss.

Generation54.3tok/s
Prompt processing1046tok/s
Max context10Ktokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 8B on a RTX 3070

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 15.0 GB 16.8 GB Won't fit 3.53
INT8 / W8A8 8.50 7.7 GB 9.5 GB Won't fit 10.5
Q8_0 (GGUF) 8.50 7.7 GB 9.5 GB Won't fit 10.5
FP8 (E4M3) 8.00 7.5 GB 9.3 GB Won't fit 10.8
Q6_K 6.56 6.1 GB 8.0 GB Won't fit 18.2
AWQ 4-bit 4.25 5.4 GB 7.2 GB Won't fit 4K 26.9
GPTQ 4-bit 4.25 5.4 GB 7.2 GB Won't fit 4K 26.9
MXFP4 4.25 5.4 GB 7.2 GB Won't fit 4K 26.9
Q5_K_M 5.67 5.3 GB 7.1 GB Won't fit 5K 30.7
Q5_K_S 5.52 5.2 GB 7.0 GB Won't fit 6K 35.8
Q4_K_M 4.85 4.6 GB 6.4 GB Fits, but tight 10K 54.3
Q4_K_S 4.58 4.4 GB 6.2 GB Fits, but tight 12K 56.7
Q4_0 4.55 4.4 GB 6.2 GB Fits, but tight 12K 57.0
IQ4_XS 4.25 4.1 GB 5.9 GB Runs comfortably 14K 59.9
Q3_K_M 3.91 3.8 GB 5.7 GB Runs comfortably 16K 63.7
IQ3_M 3.70 3.7 GB 5.5 GB Runs comfortably 18K 66.2
IQ3_XXS 3.06 3.2 GB 5.0 GB Runs comfortably 22K 75.5
Q2_K 2.63 2.8 GB 4.6 GB Runs comfortably 25K 83.4
IQ2_XXS 2.06 2.3 GB 4.2 GB Runs comfortably 28K 96.7
IQ1_M 1.75 2.1 GB 3.9 GB Runs comfortably 30K 106

Models on a RTX 3070

Q4_K_M at 8K context. Bandwidth sets the speed; capacity sets the ceiling.

ModelParamsWeightsVerdictMax ctxtok/s
DeepSeek-R1 671B-A37B 671B 379.0 GB Won't fit 1.87
Qwen3 235B-A22B 235.1B 132.8 GB Won't fit 3.02
gpt-oss 120B-A5.1B 116.8B 66.0 GB Won't fit 13.0
Llama 4 Scout 109B-A17B 109B 61.7 GB Won't fit 4.06
Qwen2.5 72B 72.7B 41.2 GB Won't fit 1.00
Llama 3.1 70B 70.5B 40.0 GB Won't fit 1.03
Mixtral 8x7B 46.7B 26.4 GB Won't fit 5.87
Command R 35B 35.0B 20.1 GB Won't fit 1.59
Yi-1.5 34B 34.4B 19.5 GB Won't fit 2.36
Qwen3 32B 32.8B 18.6 GB Won't fit 2.46
Qwen2.5-Coder 32B 32.8B 18.6 GB Won't fit 2.46
DeepSeek-R1-Distill-Qwen 32B 32.8B 18.6 GB Won't fit 2.46
Qwen3 30B-A3B 30.5B 17.3 GB Won't fit 21.3
Gemma 3 27B 27.4B 15.6 GB Won't fit 3.32
Mistral Small 24B 23.6B 13.4 GB Won't fit 3.94
gpt-oss 20B-A3.6B 20.9B 11.9 GB Won't fit 28.1
Qwen3 14B 14.8B 8.5 GB Won't fit 7.99
Phi-4 14B 14.7B 8.4 GB Won't fit 7.62
Mistral NeMo 12B 12.3B 7.0 GB Won't fit 11.8
Gemma 3 12B 12.2B 7.0 GB Won't fit 14.1
GLM-4 9B 9.4B 5.4 GB Fits, but tight 13K 49.9
Qwen3 8B 8.2B 4.7 GB Fits, but tight 8K 52.5
Llama 3.1 8B 8.0B 4.6 GB Fits, but tight 10K 54.3
DeepSeek-R1-Distill-Llama 8B 8.0B 4.6 GB Fits, but tight 10K 54.3
Qwen2.5 7B 7.6B 4.4 GB Runs comfortably 28K 60.3
Mistral 7B v0.3 7.3B 4.1 GB Runs comfortably 14K 60.1
Gemma 3 4B 4.3B 2.5 GB Runs comfortably 128K 102
Qwen3 4B 4.0B 2.3 GB Runs comfortably 26K 96.0
Llama 3.2 3B 3.2B 1.8 GB Runs comfortably 37K 121
Llama 3.2 1B 1.2B 0.7 GB Runs comfortably 128K 322

Direct answers