NVIDIA · Ada

What runs on a L40S

48 GB at 864 GB/s. The largest model that fits at Q4_K_M is Llama 3.1 70B, at roughly 12.5 tokens per second.

Runs comfortably

6.4 GB of 44.3 GB · 15%
044 GB
Weights 4.6 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

Llama 3.1 8B at Q4_K_M leaves 37.9 GB spare on a L40S. There is room to raise the context length or move up a quantisation level.

Generation99.4tok/s
Prompt processing4733tok/s
Max context128Ktokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 8B on a L40S

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 15.0 GB 16.8 GB Runs comfortably 128K 33.3
INT8 / W8A8 8.50 7.7 GB 9.5 GB Runs comfortably 128K 62.4
Q8_0 (GGUF) 8.50 7.7 GB 9.5 GB Runs comfortably 128K 62.4
FP8 (E4M3) 8.00 7.5 GB 9.3 GB Runs comfortably 128K 64.1
Q6_K 6.56 6.1 GB 8.0 GB Runs comfortably 128K 77.0
AWQ 4-bit 4.25 5.4 GB 7.2 GB Runs comfortably 128K 86.2
GPTQ 4-bit 4.25 5.4 GB 7.2 GB Runs comfortably 128K 86.2
MXFP4 4.25 5.4 GB 7.2 GB Runs comfortably 128K 86.2
Q5_K_M 5.67 5.3 GB 7.1 GB Runs comfortably 128K 87.8
Q5_K_S 5.52 5.2 GB 7.0 GB Runs comfortably 128K 90.0
Q4_K_M 4.85 4.6 GB 6.4 GB Runs comfortably 128K 99.4
Q4_K_S 4.58 4.4 GB 6.2 GB Runs comfortably 128K 104
Q4_0 4.55 4.4 GB 6.2 GB Runs comfortably 128K 104
IQ4_XS 4.25 4.1 GB 5.9 GB Runs comfortably 128K 110
Q3_K_M 3.91 3.8 GB 5.7 GB Runs comfortably 128K 116
IQ3_M 3.70 3.7 GB 5.5 GB Runs comfortably 128K 121
IQ3_XXS 3.06 3.2 GB 5.0 GB Runs comfortably 128K 138
Q2_K 2.63 2.8 GB 4.6 GB Runs comfortably 128K 152
IQ2_XXS 2.06 2.3 GB 4.2 GB Runs comfortably 128K 176
IQ1_M 1.75 2.1 GB 3.9 GB Runs comfortably 128K 192

Models on a L40S

Q4_K_M at 8K context. Bandwidth sets the speed; capacity sets the ceiling.

ModelParamsWeightsVerdictMax ctxtok/s
DeepSeek-R1 671B-A37B 671B 379.0 GB Won't fit 1.97
Qwen3 235B-A22B 235.1B 132.8 GB Won't fit 3.90
gpt-oss 120B-A5.1B 116.8B 66.0 GB Won't fit 27.8
Llama 4 Scout 109B-A17B 109B 61.7 GB Won't fit 9.40
Qwen2.5 72B 72.7B 41.2 GB Won't fit 7K 10.5
Llama 3.1 70B 70.5B 40.0 GB Fits, but tight 11K 12.5
Mixtral 8x7B 46.7B 26.4 GB Runs comfortably 32K 55.1
Command R 35B 35.0B 20.1 GB Runs comfortably 19K 20.6
Yi-1.5 34B 34.4B 19.5 GB Runs comfortably 32K 25.1
Qwen3 32B 32.8B 18.6 GB Runs comfortably 99K 26.2
Qwen2.5-Coder 32B 32.8B 18.6 GB Runs comfortably 99K 26.1
DeepSeek-R1-Distill-Qwen 32B 32.8B 18.6 GB Runs comfortably 99K 26.1
Qwen3 30B-A3B 30.5B 17.3 GB Runs comfortably 128K 109
Gemma 3 27B 27.4B 15.6 GB Runs comfortably 128K 31.4
Mistral Small 24B 23.6B 13.4 GB Runs comfortably 32K 36.6
gpt-oss 20B-A3.6B 20.9B 11.9 GB Runs comfortably 128K 166
Qwen3 14B 14.8B 8.5 GB Runs comfortably 128K 56.3
Phi-4 14B 14.7B 8.4 GB Runs comfortably 16K 55.8
Mistral NeMo 12B 12.3B 7.0 GB Runs comfortably 128K 66.7
Gemma 3 12B 12.2B 7.0 GB Runs comfortably 128K 67.5
GLM-4 9B 9.4B 5.4 GB Runs comfortably 128K 91.2
Qwen3 8B 8.2B 4.7 GB Runs comfortably 128K 96.1
Llama 3.1 8B 8.0B 4.6 GB Runs comfortably 128K 99.4
DeepSeek-R1-Distill-Llama 8B 8.0B 4.6 GB Runs comfortably 128K 99.4
Qwen2.5 7B 7.6B 4.4 GB Runs comfortably 128K 110
Mistral 7B v0.3 7.3B 4.1 GB Runs comfortably 32K 110
Gemma 3 4B 4.3B 2.5 GB Runs comfortably 128K 186
Qwen3 4B 4.0B 2.3 GB Runs comfortably 128K 174
Llama 3.2 3B 3.2B 1.8 GB Runs comfortably 128K 219
Llama 3.2 1B 1.2B 0.7 GB Runs comfortably 128K 579

Direct answers