Head to head
NVIDIA DGX Spark (GB10) vs A100 80GB
The NVIDIA DGX Spark (GB10) holds more — 128 GB against 80 GB — which decides what you can load at all. The A100 80GB has more bandwidth at 2039 GB/s, which decides how fast tokens come out once a model fits.
| Model (Q4_K_M, 8K) | NVIDIA DGX Spark (GB10) | tok/s | A100 80GB | tok/s |
|---|---|---|---|---|
| Llama 3.1 8B | Runs comfortably | 35.1 | Runs comfortably | 238 |
| Llama 3.1 70B | Runs comfortably | 4.37 | Runs comfortably | 30.5 |
| Llama 3.2 3B | Runs comfortably | 78.1 | Runs comfortably | 512 |
| Llama 3.2 1B | Runs comfortably | 209 | Runs comfortably | 1319 |
| Llama 4 Scout 109B-A17B | Runs comfortably | 14.8 | Runs comfortably | 100 |
| Qwen3 8B | Runs comfortably | 33.9 | Runs comfortably | 230 |
| Qwen3 14B | Runs comfortably | 19.8 | Runs comfortably | 136 |
| Qwen3 32B | Runs comfortably | 9.17 | Runs comfortably | 63.5 |
| Qwen3 4B | Runs comfortably | 62.2 | Runs comfortably | 407 |
| Qwen3 30B-A3B | Runs comfortably | 53.5 | Runs comfortably | 163 |
| Qwen3 235B-A22B | Won't fit | 6.20 | Won't fit | 6.02 |
| Qwen2.5-Coder 32B | Runs comfortably | 9.17 | Runs comfortably | 63.5 |
| Qwen2.5 7B | Runs comfortably | 38.9 | Runs comfortably | 265 |
| Qwen2.5 72B | Runs comfortably | 4.24 | Runs comfortably | 29.6 |
| Gemma 3 4B | Runs comfortably | 66.4 | Runs comfortably | 434 |
| Gemma 3 12B | Runs comfortably | 23.8 | Runs comfortably | 162 |
Specifications
| NVIDIA DGX Spark (GB10) | A100 80GB | |
|---|---|---|
| Memory | 128 GB | 80 GB |
| Bandwidth | 273 GB/s | 2039 GB/s |
| FP16 compute | 125 TFLOPS | 312 TFLOPS |
| Launch price | $3,999 | $15,000 |
| Architecture | Blackwell | Ampere |