Models · generated from the repo

Every model goinfer runs, and how sure we are about each one.

Each family is checked against HuggingFace's own implementation, and the result is printed next to it. A few checkpoints are vetted end to end: pinned, sized, fitted and timed. Everything else says what hasn't been shown yet.

40
model families
33
checked against the released model
6
checked against a small test model only
1
with no check of its own
6
vetted checkpoints
Fit and speed for

Vetted checkpoints: the supported tier

These are the checkpoints the project supports. Each is pinned by checksum and comes from a family checked against its released model, the strongest of the checks, with a fit verdict and whatever speed was measured. Other checkpoints of a family run through the same code with less shown for them. If you don't know where to start, start here.

Qwen2.5-Coder 0.5B

q4_k_m · 0.49 GB
fitsfitsfits

Code completion and small edits

Against the released model
171.2 tok/s 1.18× Ollama367.9 tok/s 1.40× Ollama52 tok/s 0.91× Ollamano speeds for a machine we haven't measured

Qwen2.5-Coder 1.5B

q4_k_m · 1.1 GB
fitsfitsfits

Qwen2 and Qwen2.5, including the Coder models most of goinfer's numbers are measured on.

Against the released model
89.3 tok/s 1.06× Ollama270.7 tok/s 1.47× Ollama24.3 tok/s level with Ollamano speeds for a machine we haven't measured

Phi-3 mini 4k

q4 · 2.4 GB
fits on the CPUfits on the CPUfits

General instruction following at a size that still loads in seconds

Against the released model
not measured on this machine112.6 tok/s 0.90× Ollamanot measured on this machineno speeds for a machine we haven't measured

Granite 4.0-H Tiny

q8_0 · 7.4 GB
fits on the CPUfits on the CPUfits

A Mamba-2/attention hybrid MoE, if you want to exercise that path

Against the released model
not measured on this machinenot measured on this machinenot measured on this machineno speeds for a machine we haven't measured

gpt-oss 20B

mxfp4 · 12.1 GB
tight pages experts from diskfits experts streamed to the cardfits

The 20-35B-class MoE this project actually validates and measures

Against the released model
not measured on this machinenot measured on this machinenot measured on this machineno speeds for a machine we haven't measured

Gemma 4 26B-A4B

q4_0 · 14.4 GB
tight pages experts from diskfits experts streamed to the cardfits

The 20-35B-class MoE this project has the most measurements on

Against the released model
not measured on this machine39.3 tok/s 1.76× Ollama, graded ambiguousnot measured on this machineno speeds for a machine we haven't measured

All families

Sorted by how strong the check is, strongest first. There's no "popular" sort: goinfer doesn't count downloads.

For
Checked
Needs
Showing 40 of 40 familiesWhat the checks mean

Embeddings

Separate from the families above: they are encoders, not chat models.

goinfer serve --embed-model serves /v1/embeddings, on its own or beside a chat model. Two encoders work today. Qwen3-Embedding-0.6B (a GGUF file) matches the sentence-transformers reference to cosine 1.0000000 on five test inputs, checked 2026-07-20. CodeRankEmbed (a HuggingFace directory) is served and timed against Ollama's nomic-embed-text, but no comparison with a reference is recorded here. Details in the server docs.

What the checks mean

Every family gets one of these. They are listed strongest first.

Against the released model

The family's output is compared with HuggingFace's own implementation on a real checkpoint. Two numbers are printed: how often the next token is the same, and the closest the raw scores get at their worst position (cosine, where 1.00000 is identical).

Against a small test model

No released model was small enough to run on the hardware available, so the family is checked against a small test model built from the same wiring. A weaker claim, and it is labelled that way.

Shares another family's check

The family runs another family's code, so that family's numbers stand in for it. It has none of its own.

Not recorded

No check is on file for this family yet.

More in what "parity-gated" means. A family registered as experimental can change, or go, before v1.0.

Coming from Ollama?

Of the 60 most-pulled models in Ollama's library on 2026-10-02, goinfer runs the families behind 93.3% of the pulls as Ollama ships them, and another 1.0% as text only.

Pull counts are cumulative, so older models rank higher than their current use; this is a guide to what loads, not a measure of what people run.

#Ollama taggoinferFamilyNote
1llama3.1Runs Llama
2deepseek-r1Runs Qwen2 / Qwen2.5, Qwen3, Llama, DeepSeek-V3 distills plus the 671B
3nomic-embed-textRuns Embeddings
4llama3.2Runs Llama
5qwen2.5Runs Qwen2 / Qwen2.5
6gemma3Runs Gemma 3
7qwen3Runs Qwen3
8gemma2Runs Gemma 2
9mistralRuns Mistral
10gemma4Runs Gemma 4 images yes; audio input not yet
11llama3Runs Llama
12qwen2.5-coderRuns Qwen2 / Qwen2.5
13qwen3.5Runs Qwen3.8, Qwen3.5-MoE images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text
14phi3Runs Phi-3 / Phi-4
15mxbai-embed-largeRuns Embeddings
16llavaNot supported yet
17gpt-ossRuns gpt-oss
18qwen3-coderRuns Qwen3-MoE
19gemmaRuns Gemma Gemma 1
20qwenNot supported yet Qwen 1
The other 40
#Ollama taggoinferFamilyNote
21glm-ocrRuns GLM-OCR images from the safetensors checkpoint only (no GGUF); the vision tower runs on the CPU; tool calls are not rendered; parity tier experimental
22phi4Runs Phi-3 / Phi-4 Phi-4 uses the phi3 architecture
23llama2Runs Llama
24bge-m3Runs Embeddings
25qwen3.6Runs Qwen3.8, Qwen3.5-MoE images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text
26codellamaRuns Llama
27qwen3-vlRuns as text only Qwen3-VL text side experimental; safetensors only
28qwen2Runs Qwen2 / Qwen2.5
29tinyllamaRuns Llama
30mistral-nemoRuns Mistral
31minicpm-vNot supported yet
32llama3.2-visionNot supported yet cross-attention (mllama)
33qwen2.5vlRuns Qwen2.5-VL safetensors only
34deepseek-coderRuns Llama
35llama3.3Runs Llama
36qwen3-embeddingRuns Qwen3
37dolphin3Runs Llama
38smollm2Runs Llama
39deepseek-v3Runs DeepSeek-V3
40olmo2Not supported yet goinfer has olmo3, not olmo2
41all-minilmRuns Embeddings
42codegemmaRuns Gemma the Gemma 1 architecture
43deepseek-coder-v2Runs DeepSeek-V2
44mistral-smallRuns Mistral
45snowflake-arctic-embedRuns Embeddings
46orca-miniRuns Llama
47granite3.1-moeNot supported yet granitemoe; goinfer has granite and granitemoehybrid
48qwen3.8Runs Qwen3.8, Qwen3.5-MoE images from safetensors, dense sizes only; Ollama's GGUF and the MoE sizes run as text
49starcoder2Not supported yet
50nemotron-3-superNot checked yet Nemotron-H family matches; the Super MoE variant is unchecked
51mixtralRuns Mixtral parity tier experimental
52llama2-uncensoredRuns Llama
53falcon3Runs Llama
54translategemmaRuns Gemma 3
55mistral-small3.2Runs as text only Ministral 3 vision tower ignored; safetensors only
56minimax-m2.7Not supported yet minimax_m2
57llava-llama3Not supported yet
58embeddinggemmaRuns Gemma 3
59qwqRuns Qwen2 / Qwen2.5
60gemma3nNot supported yet

Ollama's library page, read 2026-10-02. The statuses and the figures come from the snapshot.