Mistral on NVIDIA Ampere

Can I run Mixtral 8x7B on an RTX 3080 10GB?

Not at Q4_K_M — it needs 28.2 GB against 8.6 GB available. You would need 4 of these cards.

Won't fit

28.2 GB of 8.6 GB · 328%
029 GB
Weights 26.4 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB
Over the limit 19.6 GB

Short by 19.6 GB. You can run it with 8 of 32 layers on the RTX 3080 10GB and the rest in system RAM, at roughly 6.59 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation6.59tok/s
Prompt processing977tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Mixtral 8x7B on a RTX 3080 10GB

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 87.0 GB 88.8 GB Won't fit 1.72
INT8 / W8A8 8.50 46.2 GB 48.0 GB Won't fit 3.39
Q8_0 (GGUF) 8.50 46.2 GB 48.0 GB Won't fit 3.39
FP8 (E4M3) 8.00 43.5 GB 45.3 GB Won't fit 3.59
Q6_K 6.56 35.7 GB 37.5 GB Won't fit 4.63
Q5_K_M 5.67 30.8 GB 32.6 GB Won't fit 5.50
Q5_K_S 5.52 30.0 GB 31.8 GB Won't fit 5.64
Q4_K_M 4.85 26.4 GB 28.2 GB Won't fit 6.59
Q4_K_S 4.58 24.9 GB 26.7 GB Won't fit 6.95
Q4_0 4.55 24.8 GB 26.6 GB Won't fit 6.99
AWQ 4-bit 4.25 23.5 GB 25.3 GB Won't fit 7.63
GPTQ 4-bit 4.25 23.5 GB 25.3 GB Won't fit 7.63
MXFP4 4.25 23.5 GB 25.3 GB Won't fit 7.63
IQ4_XS 4.25 23.1 GB 25.0 GB Won't fit 7.73
Q3_K_M 3.91 21.3 GB 23.1 GB Won't fit 8.66
IQ3_M 3.70 20.2 GB 22.0 GB Won't fit 9.11
IQ3_XXS 3.06 16.7 GB 18.5 GB Won't fit 11.7
Q2_K 2.63 14.4 GB 16.2 GB Won't fit 15.3
IQ2_XXS 2.06 11.3 GB 13.1 GB Won't fit 23.4
IQ1_M 1.75 9.6 GB 11.4 GB Won't fit 32.6

Also worth checking