Models / Gemma 4generated from capability-matrix.json

Gemma 4

Google's Gemma 4: dense models, the small E-models, and the 26B-A4B mixture-of-experts. Reads images.

chatimagesexperts: dense ‖ sparse, no-sharedsafetensors, GGUF

Gemma 4 26B-A4B

Get it

tight pages experts from disk on MacBook Pro, M1 Pro, 16 GBfits experts streamed to the card on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull gemma-4-26b-a4b
repo
google/gemma-4-26B-A4B-it-qat-q4_0-gguf
file
gemma-4-26B_q4_0-it.gguf
quant
q4_0
size
14.4 GB
sha256
3eca3b8f6d7baf21…

Good for, and what it needs

The 20-35B-class MoE this project has the most measurements on (docs/benchmarks.md §B4/§B4.1): a 26B-A4B that does not fit an 8 GB card, kept fully GPU-resident via host↔VRAM expert streaming (the C′ cache, -moe-cache-experts) rather than CPU-offloaded.

~11.4 GB of int4 experts; does not fit an 8 GB card resident. Measured on an RTX 2070 SUPER (driver 595.91.07): 39.3 tok/s at ctx 2048 (median of three runs, 37.3–39.4), every expert kept on the GPU through host→VRAM streaming with the DMA overlap (docs/measurements/peer-sweep-2026-09-29.md cell c), at ~20.0 GiB peak host RSS — capacity-bound (PCIe host→VRAM streaming), not a kernel or MoE deficiency

Tool calling

not yet measured

Decisions · /v1/systemone

Label scoring, unmeasured.

What hasn't been shown

  • Tool calling hasn't been measured on Gemma 4 26B-A4B.
  • No graded speed on the M1 Pro 16 GB.
  • No graded speed on the Ryzen 7 3700X.
  • Never measured on a machine like yours, so no speed is shown.

How sure we are

Against the released model
100.0% picks the same next token as HuggingFace
0.98972 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

full-oracle 100.0%/0.98972 · what parity-gated means

Measured speed · decode, tokens per second

your pick · MacBook Pro, M1 Pro, 16 GBMetal
not measured
your pick · RTX 2070 SUPER, 8 GBCUDA · 2026-09-29
goinfer
39.3 Ollama
22.3
1.76× Ollama, graded ambiguous Every pair was above 1.67×, but goinfer's own runs spread 5.3%, over the 5% cap set before the run, so it is not graded ahead. goinfer keeps every expert on the card; Ollama offloads to the CPU. Different checkpoints, so this compares approaches, not kernels. Context capped at 2,048.
your pick · Ryzen 7 3700X, CPU onlyCPU
not measured

Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.

Architecture, for the curious
model_type
gemma4, gemma4_text, gemma4_unified_text
design
softmax-GQA
experts
dense ‖ sparse, no-shared
attention window
interleave
QK-norm
yes
RoPE
dual-base
norm
RMSNorm, sandwich
activation
GeGLU
tied head
yes
modality
text (+ vision tower)
GPU-resident
eligible
registry description: Google Gemma 4 dense + E-models (per-layer attention deltas, PLE)