Models · generated from the repo
Every model goinfer runs, and how sure we are about each one.
Each family is checked against HuggingFace's own implementation, and the result is printed next to it. A few checkpoints are vetted end to end: pinned, sized, fitted and timed. Everything else says what hasn't been shown yet.
Vetted checkpoints: the supported tier
These are the checkpoints the project supports. Each is pinned by checksum and comes from a family checked against its released model, the strongest of the checks, with a fit verdict and whatever speed was measured. Other checkpoints of a family run through the same code with less shown for them. If you don't know where to start, start here.
Qwen2.5-Coder 0.5B
q4_k_m · 0.49 GBCode completion and small edits
Against the released modelQwen2.5-Coder 1.5B
q4_k_m · 1.1 GBQwen2 and Qwen2.5, including the Coder models most of goinfer's numbers are measured on.
Against the released modelPhi-3 mini 4k
q4 · 2.4 GBGeneral instruction following at a size that still loads in seconds
Against the released modelGranite 4.0-H Tiny
q8_0 · 7.4 GBA Mamba-2/attention hybrid MoE, if you want to exercise that path
Against the released modelgpt-oss 20B
mxfp4 · 12.1 GBThe 20-35B-class MoE this project actually validates and measures
Against the released modelGemma 4 26B-A4B
q4_0 · 14.4 GBThe 20-35B-class MoE this project has the most measurements on
Against the released modelAll families
Sorted by how strong the check is, strongest first. There's no "popular" sort: goinfer doesn't count downloads.
-
Command-RCohere's Command-R and Aya.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Command-R7BCohere's Command-R7B and Command-A.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
GemmaGoogle's first Gemma, 2B and 7B, and CodeGemma.chatcodeAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Gemma 2Google's Gemma 2, 2B to 27B.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointCPU onlysafetensors, GGUF
-
GPT-2GPT-2 and GPT-NeoX. Old and small, kept as a reference.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Granite 4.2IBM's Granite 4.2 dense, 3B to 30B.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
InternLM2Shanghai AI Lab's InternLM2.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
LFM2.5Liquid AI's LFM2 and LFM2.5: short convolutions mixed with attention.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetCPU onlysafetensors
-
LlamaMeta's Llama 2 and 3, plus InternLM3. The one family that also loads GPTQ and AWQ.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF, GPTQ, AWQ
-
Ministral 3Ministral 3, 3B to 14B. Text only: the vision tower is skipped.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
MistralMistral-style dense models with a sliding attention window.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Olmo 3Ai2's Olmo 3, 7B and 32B.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Olmo HybridAi2's 7B hybrid: mostly linear-attention layers, a few full-attention ones, no position encoding at all.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Phi-3 / Phi-4Microsoft's Phi-3 and Phi-4.chatAgainst the released modelfull-oracle 100.0%/1.00000Phi-3 mini 4kvetted, one commandCPU onlysafetensors, GGUF
-
Qwen2 / Qwen2.5Qwen2 and Qwen2.5, including the Coder models most of goinfer's numbers are measured on.chatcodeAgainst the released modelfull-oracle 100.0%/1.000002 checkpointsvetted, one commandGPU or CPUsafetensors, GGUF
-
Qwen3Alibaba's Qwen3 dense models.chatAgainst the released modelfull-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
SmolLM3Hugging Face's SmolLM3, 3B.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Spark-X2.5XHToken's Spark-X2.5, 1.7B and 4B.chatAgainst the released modelfull-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetCPU onlysafetensors
-
Qwen3.8Qwen3.8 dense: the same hybrid attention, without the experts. Reads images (verified on Qwen3.5-0.8B).chatimagesAgainst the released modelreal-oracle 100.0%/0.99989Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Gemma 3Google's Gemma 3, from 270M to 27B. The 4B and larger read images (verified on the 4B).chatimagesAgainst the released modelfull-oracle 100.0%/0.99972Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Mellum2JetBrains' Mellum2, built for code completion.codeexpertsAgainst the released modelreal-oracle 100.0%/0.99969Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Qwen2-MoEQwen1.5 and Qwen2 mixture-of-experts.chatexpertsAgainst the released modelreal-oracle 100.0%/0.99954Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
DeepSeek-V3DeepSeek-V3: the same design at frontier scale. Far larger than any machine on this site.chatexpertsAgainst the released modelreal-oracle 100.0%/0.99951Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Qwen2.5-VLQwen2.5-VL. Reads images.chatimagesAgainst the released modelfull-oracle 100.0%/0.99946Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
DeepSeek-V2DeepSeek-V2 and V2-Lite: a compressed attention cache plus experts.chatexpertsAgainst the released modelreal-oracle 100.0%/0.99924Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Lagunapoolside's Laguna coding models.codeexpertsAgainst the released modelreal-oracle 100.0%/0.99884Pull by repoGGUF · no vetted checkpointCPU onlysafetensors, GGUF
-
gpt-ossOpenAI's open-weight gpt-oss, 20B and 120B.chatexpertsAgainst the released modelreal-oracle 100.0%/0.99843gpt-oss 20Bvetted, one commandGPU or CPUsafetensors, GGUF
-
Qwen3-MoEQwen3-30B-A3B and Qwen3-Coder-30B-A3B.chatcodeexpertsAgainst the released modelreal-oracle 100.0%/0.99834Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Nemotron-HNVIDIA's Nemotron-H and Nemotron 3 Nano: a Mamba-2 hybrid.chatAgainst the released modelreal-oracle 100.0%/0.99574Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Granite-4.0-HIBM's Granite 4.0-H: Mamba-2 layers plus attention and experts.chatexpertsAgainst the released modelreal-oracle 100.0%/0.99566Granite 4.0-H Tinyvetted, one commandCPU onlysafetensors, GGUF
-
Gemma 4Google's Gemma 4: dense models, the small E-models, and the 26B-A4B mixture-of-experts. Reads images.chatimagesexpertsAgainst the released modelfull-oracle 100.0%/0.98972Gemma 4 26B-A4Bvetted, one commandGPU or CPUsafetensors, GGUF
-
Qwen3-NextQwen's 80B mixture-of-experts with about 3B active per token, on hybrid attention.chatexpertsAgainst the released modelreal-oracle 100.0%/0.98931Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Qwen3.5-MoEQwen3.5 and 3.6 mixture-of-experts on hybrid attention.chatexpertsAgainst the released modelfull-oracle 77.5%/0.99069Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
GLM-4.5/4.6Zhipu's GLM-4.5 and 4.6 mixture-of-experts.chatexpertsexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
GLM-OCRZhipu's GLM-OCR document model. Reads an image of a page: text, a table or a formula out, or, given a schema, JSON that matches it (goinfer-chat --image invoice.png --schema invoice.schema.json). Verified on the real checkpoint, token-identical to transformers at f32. Images up to 6,144 image tokens, about 4.8 megapixels (the default); --vision-max-pixels lowers the cap, and the CPU vision tower's cost falls with it.chatimagesexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Ling 3.0inclusionAI's Ling 3.0, tiny and flash.chatexpertsexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetCPU onlysafetensors
-
Llama 4Meta's Llama 4 Scout and Maverick, text only.chatexpertsexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000 +coherentPull by repoGGUF · no vetted checkpointCPU onlysafetensors, GGUF
-
MixtralMistral's mixture-of-experts.chatexpertsexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
-
Qwen3-VLQwen3-VL, the text half only. It doesn't take images yet.chatexperimentalAgainst a small test modeltiny-oracle 100.0%/1.00000Download it yourselfsafetensors only · can't pull yetGPU or CPUsafetensors
-
Kimi K2Moonshot's Kimi K2 line, K2.7-Code included. Runs on DeepSeek-V3's code path.chatcodeexpertsShares another family's checkshared-path: deepseek_v3Pull by repoGGUF · no vetted checkpointGPU or CPUsafetensors, GGUF
- Nothing matches all of those. Loosen a filter.
Embeddings
Separate from the families above: they are encoders, not chat models.
goinfer serve --embed-model serves /v1/embeddings, on its own or beside a chat model. Two encoders work today. Qwen3-Embedding-0.6B (a GGUF file) matches the sentence-transformers reference to cosine 1.0000000 on five test inputs, checked 2026-07-20. CodeRankEmbed (a HuggingFace directory) is served and timed against Ollama's nomic-embed-text, but no comparison with a reference is recorded here. Details in the server docs.
What the checks mean
Every family gets one of these. They are listed strongest first.
The family's output is compared with HuggingFace's own implementation on a real checkpoint. Two numbers are printed: how often the next token is the same, and the closest the raw scores get at their worst position (cosine, where 1.00000 is identical).
No released model was small enough to run on the hardware available, so the family is checked against a small test model built from the same wiring. A weaker claim, and it is labelled that way.
The family runs another family's code, so that family's numbers stand in for it. It has none of its own.
No check is on file for this family yet.
More in what "parity-gated" means. A family registered as experimental can change, or go, before v1.0.
Coming from Ollama?
Of the 60 most-pulled models in Ollama's library on 2026-10-02, goinfer runs the families behind 93.3% of the pulls as Ollama ships them, and another 1.0% as text only.
Pull counts are cumulative, so older models rank higher than their current use; this is a guide to what loads, not a measure of what people run.
| # | Ollama tag | goinfer | Family | Note |
|---|---|---|---|---|
| 1 | llama3.1 | Runs | Llama | |
| 2 | deepseek-r1 | Runs | Qwen2 / Qwen2.5, Qwen3, Llama, DeepSeek-V3 | distills plus the 671B |
| 3 | nomic-embed-text | Runs | Embeddings | |
| 4 | llama3.2 | Runs | Llama | |
| 5 | qwen2.5 | Runs | Qwen2 / Qwen2.5 | |
| 6 | gemma3 | Runs | Gemma 3 | |
| 7 | qwen3 | Runs | Qwen3 | |
| 8 | gemma2 | Runs | Gemma 2 | |
| 9 | mistral | Runs | Mistral | |
| 10 | gemma4 | Runs | Gemma 4 | images yes; audio input not yet |
| 11 | llama3 | Runs | Llama | |
| 12 | qwen2.5-coder | Runs | Qwen2 / Qwen2.5 | |
| 13 | qwen3.5 | Runs | Qwen3.8, Qwen3.5-MoE | images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text |
| 14 | phi3 | Runs | Phi-3 / Phi-4 | |
| 15 | mxbai-embed-large | Runs | Embeddings | |
| 16 | llava | Not supported yet | ||
| 17 | gpt-oss | Runs | gpt-oss | |
| 18 | qwen3-coder | Runs | Qwen3-MoE | |
| 19 | gemma | Runs | Gemma | Gemma 1 |
| 20 | qwen | Not supported yet | Qwen 1 |
The other 40
| # | Ollama tag | goinfer | Family | Note |
|---|---|---|---|---|
| 21 | glm-ocr | Runs | GLM-OCR | images from the safetensors checkpoint only (no GGUF); the vision tower runs on the CPU; tool calls are not rendered; parity tier experimental |
| 22 | phi4 | Runs | Phi-3 / Phi-4 | Phi-4 uses the phi3 architecture |
| 23 | llama2 | Runs | Llama | |
| 24 | bge-m3 | Runs | Embeddings | |
| 25 | qwen3.6 | Runs | Qwen3.8, Qwen3.5-MoE | images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text |
| 26 | codellama | Runs | Llama | |
| 27 | qwen3-vl | Runs as text only | Qwen3-VL | text side experimental; safetensors only |
| 28 | qwen2 | Runs | Qwen2 / Qwen2.5 | |
| 29 | tinyllama | Runs | Llama | |
| 30 | mistral-nemo | Runs | Mistral | |
| 31 | minicpm-v | Not supported yet | ||
| 32 | llama3.2-vision | Not supported yet | cross-attention (mllama) | |
| 33 | qwen2.5vl | Runs | Qwen2.5-VL | safetensors only |
| 34 | deepseek-coder | Runs | Llama | |
| 35 | llama3.3 | Runs | Llama | |
| 36 | qwen3-embedding | Runs | Qwen3 | |
| 37 | dolphin3 | Runs | Llama | |
| 38 | smollm2 | Runs | Llama | |
| 39 | deepseek-v3 | Runs | DeepSeek-V3 | |
| 40 | olmo2 | Not supported yet | goinfer has olmo3, not olmo2 | |
| 41 | all-minilm | Runs | Embeddings | |
| 42 | codegemma | Runs | Gemma | the Gemma 1 architecture |
| 43 | deepseek-coder-v2 | Runs | DeepSeek-V2 | |
| 44 | mistral-small | Runs | Mistral | |
| 45 | snowflake-arctic-embed | Runs | Embeddings | |
| 46 | orca-mini | Runs | Llama | |
| 47 | granite3.1-moe | Not supported yet | granitemoe; goinfer has granite and granitemoehybrid | |
| 48 | qwen3.8 | Runs | Qwen3.8, Qwen3.5-MoE | images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text |
| 49 | starcoder2 | Not supported yet | ||
| 50 | nemotron-3-super | Not checked yet | Nemotron-H | family matches; the Super MoE variant is unchecked |
| 51 | mixtral | Runs | Mixtral | parity tier experimental |
| 52 | llama2-uncensored | Runs | Llama | |
| 53 | falcon3 | Runs | Llama | |
| 54 | translategemma | Runs | Gemma 3 | |
| 55 | mistral-small3.2 | Runs as text only | Ministral 3 | vision tower ignored; safetensors only |
| 56 | minimax-m2.7 | Not supported yet | minimax_m2 | |
| 57 | llava-llama3 | Not supported yet | ||
| 58 | embeddinggemma | Runs | Gemma 3 | |
| 59 | qwq | Runs | Qwen2 / Qwen2.5 | |
| 60 | gemma3n | Not supported yet |
Ollama's library page, read 2026-10-02. The statuses and the figures come from the snapshot.