Gemma · Gemma 3 · 27.4B parameters

Gemma 3 27B VRAM requirements

Gemma 3 27B has 62 layers and uses grouped-query attention (16 KV heads). At Q4_K_M the weights come to 15.6 GB, and the best quantisation that fits a 24 GB card is Q5_K_M.

Runs comfortably

17.5 GB of 21.8 GB · 80%
022 GB
Weights 15.6 GB
KV cache 1.1 GB
Runtime overhead 0.8 GB

Gemma 3 27B at Q4_K_M leaves 4.3 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation36.5tok/s
Prompt processing1265tok/s
Max context57Ktokens
KV per 1K tokens0GB

Every quantisation of Gemma 3 27B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 51.0 GB 53.0 GB Won't fit 1.16
INT8 / W8A8 8.50 26.8 GB 28.7 GB Won't fit 4.33
Q8_0 (GGUF) 8.50 26.8 GB 28.7 GB Won't fit 4.33
FP8 (E4M3) 8.00 25.5 GB 27.4 GB Won't fit 5.29
Q6_K 6.56 20.9 GB 22.9 GB Won't fit 14.2
Q5_K_M 5.67 18.1 GB 20.0 GB Fits, but tight 28K 31.7
Q5_K_S 5.52 17.6 GB 19.5 GB Runs comfortably 34K 32.5
Q4_K_M 4.85 15.6 GB 17.5 GB Runs comfortably 57K 36.5
AWQ 4-bit 4.25 15.5 GB 17.4 GB Runs comfortably 59K 36.7
GPTQ 4-bit 4.25 15.5 GB 17.4 GB Runs comfortably 59K 36.7
MXFP4 4.25 15.5 GB 17.4 GB Runs comfortably 59K 36.7
Q4_K_S 4.58 14.8 GB 16.7 GB Runs comfortably 67K 38.4
Q4_0 4.55 14.7 GB 16.6 GB Runs comfortably 68K 38.7
IQ4_XS 4.25 13.8 GB 15.7 GB Runs comfortably 79K 41.0
Q3_K_M 3.91 12.7 GB 14.7 GB Runs comfortably 91K 44.1
IQ3_M 3.70 12.1 GB 14.0 GB Runs comfortably 98K 46.3
IQ3_XXS 3.06 10.2 GB 12.1 GB Runs comfortably 121K 54.3
Q2_K 2.63 8.9 GB 10.8 GB Runs comfortably 128K 61.5
IQ2_XXS 2.06 7.1 GB 9.1 GB Runs comfortably 128K 74.6
IQ1_M 1.75 6.2 GB 8.1 GB Runs comfortably 128K 84.4

Gemma 3 27B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 136
A100 80GB 80 2039 Runs comfortably 128K 76.0
RTX 5090 32 1792 Runs comfortably 128K 70.7
RTX 4090 24 1008 Runs comfortably 57K 36.5
RTX 3090 24 936 Runs comfortably 57K 35.4
Radeon RX 7900 XTX 24 960 Runs comfortably 57K 33.1
L40S 48 864 Runs comfortably 128K 31.4
RTX A6000 48 768 Runs comfortably 128K 29.1
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 28.7
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 21.4
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 11.0
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 10.7
RTX 5080 16 960 Won't fit 9.29
RTX 5070 Ti 16 896 Won't fit 9.17
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 8.21
RTX 4080 Super 16 736 Won't fit 7.98
RTX 4070 Ti Super 16 672 Won't fit 7.81
RTX 5060 Ti 16GB 16 448 Won't fit 7.67
RTX 4060 Ti 16GB 16 288 Won't fit 5.89
RTX 5070 12 672 Won't fit 5.12
RTX 4070 Super 12 504 Won't fit 4.49
RTX 4070 12 504 Won't fit 4.49
RTX 3060 12GB 12 360 Won't fit 4.45
RTX 3080 10GB 10 760 Won't fit 3.97
Arc B580 12 456 Won't fit 3.75

Architecture

Parameters27.4B
Layers62
Hidden size5376
Attention heads / KV heads32 / 16
Head dimension128
Vocabulary262,144
Trained context128K
Sliding window1024 (every 6th layer is global)
KV cache per 1K tokens0 GB
Hugging Facegoogle/gemma-3-27b-it

The Gemma family

Google's open models. Two things dominate the memory: a 262k vocabulary — on Gemma 3 1B the embedding table is about a third of the file — and sliding-window attention, where only every sixth layer of Gemma 3 sees the full context, and every second layer on Gemma 2. The cache grows far more slowly than the context length suggests.

huggingface.co/google · ai.google.dev/gemma · all 6 Gemma models

Direct answers