Gemma

Three vision-language models, from a 4B dense decoder to a 26B sparse mixture of experts.

All three Gemma entries are vision-language: an image goes in, text comes out. They are also the family that covers the size range of the matrix most evenly — a 4B dense decoder at the small end, a 12B dense one in the middle, and a 26B sparse mixture of experts that reads about 4B parameters per token.

Comparing the three is the clearest illustration of why parameter count alone does not predict speed. The dense models read everything they hold on every token; the sparse one does not, so it sits at a different point on the size-versus-rate curve than its total suggests. The benchmarks page has the figures.

Gemma 3 4B IT

The small end of the matrix. Vision plus text at 4B, fast enough that the model stops being the bottleneck.

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 4B
  • Architecture: Dense transformer + vision encoder
  • Quantization: int4 RTN, group 128
  • Base model: google/gemma-3-4b-it
  • Validation status: Validated
  • Measured in v0.5.1: 62.98 tokens/s, 1.41 s to first token, at a 2048-token prompt — all three lengths

Gemma 4 12B IT

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 12B
  • Architecture: Dense transformer + vision encoder
  • Quantization: int4 k-quant, block 32
  • Base model: google/gemma-4-12B-it
  • Validation status: Validated
  • Measured in v0.5.1: 21.85 tokens/s, 2.46 s to first token, at a 2048-token prompt — all three lengths

Gemma 4 26B-A4B IT

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 26B total, 4B active
  • Architecture: Sparse mixture of experts + vision encoder
  • Quantization: int4 weights, quantized MoE
  • Base model: google/gemma-4-26B-A4B-it
  • Validation status: Validated
  • Measured in v0.5.1: 47.21 tokens/s, 1.80 s to first token, at a 2048-token prompt — all three lengths