Llama on Apple M4

Can I run Llama 3.1 70B on an Mac Mini M4 Pro 48GB?

Not at Q4_K_M — it needs 43.1 GB against 36.0 GB available. Drop to Q3_K_M and it fits, at about 5.19 tokens per second.

Won't fit

43.1 GB of 36.0 GB · 120%
045 GB
Weights 40.0 GB
KV cache 2.5 GB
Runtime overhead 0.6 GB
Over the limit 7.1 GB

Short by 7.1 GB. You can run it with 65 of 80 layers on the Mac Mini M4 Pro 48GB and the rest in system RAM, at roughly 2.63 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation2.63tok/s
Prompt processing50.6tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 70B on a Mac Mini M4 Pro 48GB

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 131.4 GB 134.5 GB Won't fit 0.38
INT8 / W8A8 8.50 69.3 GB 72.4 GB Won't fit 0.90
Q8_0 (GGUF) 8.50 69.3 GB 72.4 GB Won't fit 0.90
FP8 (E4M3) 8.00 65.7 GB 68.8 GB Won't fit 0.99
Q6_K 6.56 53.9 GB 57.0 GB Won't fit 1.38
Q5_K_M 5.67 46.6 GB 49.7 GB Won't fit 1.85
Q5_K_S 5.52 45.3 GB 48.4 GB Won't fit 1.98
Q4_K_M 4.85 40.0 GB 43.1 GB Won't fit 2.63
Q4_K_S 4.58 37.8 GB 40.9 GB Won't fit 3.09
AWQ 4-bit 4.25 37.8 GB 40.9 GB Won't fit 3.10
GPTQ 4-bit 4.25 37.8 GB 40.9 GB Won't fit 3.10
MXFP4 4.25 37.8 GB 40.9 GB Won't fit 3.10
Q4_0 4.55 37.6 GB 40.7 GB Won't fit 3.20
IQ4_XS 4.25 35.2 GB 38.3 GB Won't fit 640 3.86
Q3_K_M 3.91 32.5 GB 35.6 GB Fits, but tight 9K 5.19
IQ3_M 3.70 30.8 GB 33.9 GB Fits, but tight 15K 5.46
IQ3_XXS 3.06 25.7 GB 28.8 GB Runs comfortably 31K 6.49
Q2_K 2.63 22.3 GB 25.4 GB Runs comfortably 42K 7.43
IQ2_XXS 2.06 17.8 GB 20.9 GB Runs comfortably 56K 9.20
IQ1_M 1.75 15.3 GB 18.4 GB Runs comfortably 64K 10.6

Also worth checking