Qwen on NVIDIA Ada

Can I run Qwen2.5 72B on an L40S?

Not at Q4_K_M — it needs 44.6 GB against 44.3 GB available. Drop to Q4_K_S and it fits, at about 12.8 tokens per second.

Won't fit

44.6 GB of 44.3 GB · 101%
046 GB
Weights 41.2 GB
KV cache 2.5 GB
Runtime overhead 0.9 GB
Over the limit 0.3 GB

Short by 0.3 GB. You can run it with 79 of 80 layers on the L40S and the rest in system RAM, at roughly 10.5 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation10.5tok/s
Prompt processing523tok/s
Max context7Ktokens
KV per 1K tokens0GB

Every quantisation of Qwen2.5 72B on a L40S

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 135.4 GB 138.8 GB Won't fit 0.39
INT8 / W8A8 8.50 71.4 GB 74.8 GB Won't fit 1.10
Q8_0 (GGUF) 8.50 71.4 GB 74.8 GB Won't fit 1.10
FP8 (E4M3) 8.00 67.7 GB 71.1 GB Won't fit 1.25
Q6_K 6.56 55.5 GB 58.9 GB Won't fit 2.05
Q5_K_M 5.67 48.0 GB 51.4 GB Won't fit 3.65
Q5_K_S 5.52 46.7 GB 50.1 GB Won't fit 4.20
Q4_K_M 4.85 41.2 GB 44.6 GB Won't fit 7K 10.5
AWQ 4-bit 4.25 39.4 GB 42.8 GB Fits, but tight 13K 12.7
GPTQ 4-bit 4.25 39.4 GB 42.8 GB Fits, but tight 13K 12.7
MXFP4 4.25 39.4 GB 42.8 GB Fits, but tight 13K 12.7
Q4_K_S 4.58 39.0 GB 42.4 GB Fits, but tight 14K 12.8
Q4_0 4.55 38.8 GB 42.2 GB Fits, but tight 15K 12.9
IQ4_XS 4.25 36.3 GB 39.7 GB Runs comfortably 23K 13.7
Q3_K_M 3.91 33.6 GB 36.9 GB Runs comfortably 32K 14.8
IQ3_M 3.70 31.8 GB 35.2 GB Runs comfortably 37K 15.5
IQ3_XXS 3.06 26.6 GB 30.0 GB Runs comfortably 54K 18.4
Q2_K 2.63 23.1 GB 26.5 GB Runs comfortably 65K 21.1
IQ2_XXS 2.06 18.4 GB 21.8 GB Runs comfortably 80K 26.0
IQ1_M 1.75 15.9 GB 19.3 GB Runs comfortably 88K 29.8

Also worth checking