Models / Gemma 3generated from capability-matrix.json

Gemma 3

Google's Gemma 3, from 270M to 27B. The 4B and larger read images (verified on the 4B).

chatimagessafetensors, GGUF

Get it

No vetted checkpoint yet. You can pull any GGUF of this family by repo and quant. goinfer checks it fits before downloading, but nobody here has timed it.

goinfer-chat pull <owner>/<repo>:<quant>

What hasn't been shown

  • No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.

How sure we are

Against the released model
100.0% picks the same next token as HuggingFace
0.99972 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

full-oracle 100.0%/0.99972 · what parity-gated means

Measured speed · decode, tokens per second

not measured — no vetted checkpoint to measure

Architecture, for the curious
model_type
gemma3, gemma3_text
design
softmax-GQA
experts
dense
attention window
interleave
QK-norm
yes
RoPE
dual-base
norm
RMSNorm, sandwich
activation
GeGLU
tied head
yes
modality
text (+ vision via VL text_config)
GPU-resident
eligible
registry description: Google Gemma 3 dense (270M/1B/4B/12B/27B)