Llama · 8.0B parameters
Llama 3.1 8B VRAM requirements
Llama 3.1 8B has 32 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 4.6 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.
Runs comfortably
6.4 GB of 21.8 GB · 30%022 GB
Weights
4.6 GB
KV cache
1.0 GB
Runtime overhead
0.8 GB
Llama 3.1 8B at Q4_K_M leaves 15.3 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.
Generation116tok/s
Prompt processing4315tok/s
Max context128Ktokens
KV per 1K tokens0GB
Every quantisation of Llama 3.1 8B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 15.0 GB | 16.8 GB | Runs comfortably | 48K | 38.8 |
| INT8 / W8A8 | 8.50 | 7.7 GB | 9.5 GB | Runs comfortably | 106K | 72.6 |
| Q8_0 (GGUF) | 8.50 | 7.7 GB | 9.5 GB | Runs comfortably | 106K | 72.6 |
| FP8 (E4M3) | 8.00 | 7.5 GB | 9.3 GB | Runs comfortably | 108K | 74.7 |
| Q6_K | 6.56 | 6.1 GB | 8.0 GB | Runs comfortably | 118K | 89.6 |
| AWQ 4-bit | 4.25 | 5.4 GB | 7.2 GB | Runs comfortably | 124K | 100 |
| GPTQ 4-bit | 4.25 | 5.4 GB | 7.2 GB | Runs comfortably | 124K | 100 |
| MXFP4 | 4.25 | 5.4 GB | 7.2 GB | Runs comfortably | 124K | 100 |
| Q5_K_M | 5.67 | 5.3 GB | 7.1 GB | Runs comfortably | 125K | 102 |
| Q5_K_S | 5.52 | 5.2 GB | 7.0 GB | Runs comfortably | 126K | 105 |
| Q4_K_M | 4.85 | 4.6 GB | 6.4 GB | Runs comfortably | 128K | 116 |
| Q4_K_S | 4.58 | 4.4 GB | 6.2 GB | Runs comfortably | 128K | 121 |
| Q4_0 | 4.55 | 4.4 GB | 6.2 GB | Runs comfortably | 128K | 121 |
| IQ4_XS | 4.25 | 4.1 GB | 5.9 GB | Runs comfortably | 128K | 127 |
| Q3_K_M | 3.91 | 3.8 GB | 5.7 GB | Runs comfortably | 128K | 135 |
| IQ3_M | 3.70 | 3.7 GB | 5.5 GB | Runs comfortably | 128K | 141 |
| IQ3_XXS | 3.06 | 3.2 GB | 5.0 GB | Runs comfortably | 128K | 160 |
| Q2_K | 2.63 | 2.8 GB | 4.6 GB | Runs comfortably | 128K | 176 |
| IQ2_XXS | 2.06 | 2.3 GB | 4.2 GB | Runs comfortably | 128K | 204 |
| IQ1_M | 1.75 | 2.1 GB | 3.9 GB | Runs comfortably | 128K | 223 |
Llama 3.1 8B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| H100 SXM 80GB | 80 | 3350 | Runs comfortably | 128K | 422 |
| A100 80GB | 80 | 2039 | Runs comfortably | 128K | 238 |
| RTX 5090 | 32 | 1792 | Runs comfortably | 128K | 222 |
| RTX 5080 | 16 | 960 | Runs comfortably | 70K | 121 |
| RTX 4090 | 24 | 1008 | Runs comfortably | 128K | 116 |
| RTX 5070 Ti | 16 | 896 | Runs comfortably | 70K | 113 |
| RTX 3090 | 24 | 936 | Runs comfortably | 128K | 112 |
| Radeon RX 7900 XTX | 24 | 960 | Runs comfortably | 128K | 105 |
| L40S | 48 | 864 | Runs comfortably | 128K | 99.4 |
| RTX A6000 | 48 | 768 | Runs comfortably | 128K | 92.3 |
| RTX 3080 10GB | 10 | 760 | Runs comfortably | 25K | 91.4 |
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 128K | 90.9 |
| RTX 5070 | 12 | 672 | Runs comfortably | 40K | 85.4 |
| RTX 4080 Super | 16 | 736 | Runs comfortably | 70K | 84.9 |
| RTX 4070 Ti Super | 16 | 672 | Runs comfortably | 70K | 77.6 |
| Mac Studio M4 Max 128GB | 128 | 546 | Runs comfortably | 128K | 67.8 |
| RTX 4070 Super | 12 | 504 | Runs comfortably | 40K | 58.4 |
| RTX 4070 | 12 | 504 | Runs comfortably | 40K | 58.4 |
| RTX 5060 Ti 16GB | 16 | 448 | Runs comfortably | 70K | 57.3 |
| Arc B580 | 12 | 456 | Runs comfortably | 40K | 44.9 |
| RTX 3060 12GB | 12 | 360 | Runs comfortably | 40K | 43.7 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Runs comfortably | 128K | 35.1 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Runs comfortably | 128K | 34.1 |
| RTX 4060 Ti 16GB | 16 | 288 | Runs comfortably | 70K | 33.6 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Runs comfortably | 128K | 26.2 |
Architecture
| Parameters | 8.0B |
| Layers | 32 |
| Hidden size | 4096 |
| Attention heads / KV heads | 32 / 8 |
| Head dimension | 128 |
| Vocabulary | 128,256 |
| Trained context | 128K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | meta-llama/Llama-3.1-8B-Instruct |
Direct answers
Llama 3.1 8B on RTX 5090
See the verdictLlama 3.1 8B on RTX 4090
See the verdictLlama 3.1 8B on RTX 3090
See the verdictLlama 3.1 8B on RTX 5080
See the verdictLlama 3.1 8B on RTX 5070 Ti
See the verdictLlama 3.1 8B on RTX 5070
See the verdictLlama 3.1 8B on RTX 4070 Ti Super
See the verdictLlama 3.1 8B on RTX 4070 Super
See the verdictLlama 3.1 8B on RTX 4070
See the verdict
See the verdictLlama 3.1 8B on RTX 4090
See the verdictLlama 3.1 8B on RTX 3090
See the verdictLlama 3.1 8B on RTX 5080
See the verdictLlama 3.1 8B on RTX 5070 Ti
See the verdictLlama 3.1 8B on RTX 5070
See the verdictLlama 3.1 8B on RTX 4070 Ti Super
See the verdictLlama 3.1 8B on RTX 4070 Super
See the verdictLlama 3.1 8B on RTX 4070
See the verdict