Head to head
RTX 4090 vs RTX 5080
The RTX 4090 holds more — 24 GB against 16 GB — which decides what you can load at all. The RTX 4090 has more bandwidth at 1008 GB/s, which decides how fast tokens come out once a model fits.
| Model (Q4_K_M, 8K) | RTX 4090 | tok/s | RTX 5080 | tok/s |
|---|---|---|---|---|
| Llama 3.1 8B | Runs comfortably | 116 | Runs comfortably | 121 |
| Llama 3.1 70B | Won't fit | 1.60 | Won't fit | 1.35 |
| Llama 3.2 3B | Runs comfortably | 254 | Runs comfortably | 266 |
| Llama 3.2 1B | Runs comfortably | 669 | Runs comfortably | 700 |
| Llama 4 Scout 109B-A17B | Won't fit | 5.18 | Won't fit | 4.91 |
| Qwen3 8B | Runs comfortably | 112 | Runs comfortably | 117 |
| Qwen3 14B | Runs comfortably | 65.5 | Runs comfortably | 68.7 |
| Qwen3 32B | Fits, but tight | 30.5 | Won't fit | 4.98 |
| Qwen3 4B | Runs comfortably | 202 | Runs comfortably | 212 |
| Qwen3 30B-A3B | Runs comfortably | 118 | Won't fit | 46.2 |
| Qwen3 235B-A22B | Won't fit | 3.23 | Won't fit | 3.36 |
| Qwen2.5-Coder 32B | Fits, but tight | 30.5 | Won't fit | 4.98 |
| Qwen2.5 7B | Runs comfortably | 128 | Runs comfortably | 135 |
| Qwen2.5 72B | Won't fit | 1.53 | Won't fit | 1.31 |
| Gemma 3 4B | Runs comfortably | 215 | Runs comfortably | 226 |
| Gemma 3 12B | Runs comfortably | 78.5 | Runs comfortably | 82.3 |
Specifications
| RTX 4090 | RTX 5080 | |
|---|---|---|
| Memory | 24 GB | 16 GB |
| Bandwidth | 1008 GB/s | 960 GB/s |
| FP16 compute | 165 TFLOPS | 112 TFLOPS |
| Launch price | $1,599 | $999 |
| Architecture | Ada | Blackwell |