Gemma 4
Google's Gemma 4: dense models, the small E-models, and the 26B-A4B mixture-of-experts. Reads images.
Gemma 4 26B-A4B
Get it
goinfer-chat pull gemma-4-26b-a4b- repo
- google/gemma-4-26B-A4B-it-qat-q4_0-gguf
- file
- gemma-4-26B_q4_0-it.gguf
- quant
- q4_0
- size
- 14.4 GB
- sha256
- 3eca3b8f6d7baf21…
Good for, and what it needs
The 20-35B-class MoE this project has the most measurements on (docs/benchmarks.md §B4/§B4.1): a 26B-A4B that does not fit an 8 GB card, kept fully GPU-resident via host↔VRAM expert streaming (the C′ cache, -moe-cache-experts) rather than CPU-offloaded.
~11.4 GB of int4 experts; does not fit an 8 GB card resident. Measured on an RTX 2070 SUPER (driver 595.91.07): 39.3 tok/s at ctx 2048 (median of three runs, 37.3–39.4), every expert kept on the GPU through host→VRAM streaming with the DMA overlap (docs/measurements/peer-sweep-2026-09-29.md cell c), at ~20.0 GiB peak host RSS — capacity-bound (PCIe host→VRAM streaming), not a kernel or MoE deficiency
Tool calling
not yet measured
Decisions · /v1/systemone
Label scoring, unmeasured.
What hasn't been shown
- Tool calling hasn't been measured on Gemma 4 26B-A4B.
- No graded speed on the M1 Pro 16 GB.
- No graded speed on the Ryzen 7 3700X.
- Never measured on a machine like yours, so no speed is shown.
How sure we are
Against the released modelfull-oracle 100.0%/0.98972 · what parity-gated means
Measured speed · decode, tokens per second
Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.
Architecture, for the curious
- model_type
- gemma4, gemma4_text, gemma4_unified_text
- design
- softmax-GQA
- experts
- dense ‖ sparse, no-shared
- attention window
- interleave
- QK-norm
- yes
- RoPE
- dual-base
- norm
- RMSNorm, sandwich
- activation
- GeGLU
- tied head
- yes
- modality
- text (+ vision tower)
- GPU-resident
- eligible