Qwen2 / Qwen2.5
Qwen2 and Qwen2.5, including the Coder models most of goinfer's numbers are measured on.
Qwen2.5-Coder 0.5B
Get it
goinfer-chat pull qwen2.5-coder-0.5b- repo
- Qwen/Qwen2.5-Coder-0.5B-Instruct-GGUF
- file
- qwen2.5-coder-0.5b-instruct-q4_k_m.gguf
- quant
- q4_k_m
- size
- 0.49 GB
- sha256
- 1d9614638d18024d…
Good for, and what it needs
Code completion and small edits; the smallest checkpoint here that writes usable Go.
~1 GB resident at --quant int4; runs on any laptop
Tool calling
minimal schema: ok; harness-scale (12 tools): skip — too small (measured 2026-09-07, nobara-pc)
Decisions · /v1/systemone
Label scoring, unmeasured.
What hasn't been shown
- Too small for an agent's full tool list: given twelve tools, it answered instead of calling one.
- Never measured on a machine like yours, so no speed is shown.
How sure we are
Against the released modelfull-oracle 100.0%/1.00000 · what parity-gated means
Measured speed · decode, tokens per second
Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.
Qwen2.5-Coder 1.5B
Get it
goinfer-chat pull Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:q4_k_m- repo
- Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
- file
- qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
- quant
- q4_k_m
- size
- 1.1 GB
- sha256
- cc324af070c2ecbf…
Good for, and what it needs
Pinned as a release tier (pull/curated.json), not a registry checkpoint, so it carries no good-for, needs or tools fields.
Tool calling
not yet measured
Decisions · /v1/systemone
Label scoring, chat-v1, calibrated: top-1 0.2774, ECE 0.2102, over 872 out-of-distribution questions from a public decisions corpus. Measured 2026-09-28 on the RTX 2070 SUPER.
Label scoring reads the model's own probabilities for the option labels; it is not a trained decision head, and the record explains how far behind one it sits: the measurement.
What hasn't been shown
- Tool calling hasn't been measured on Qwen2.5-Coder 1.5B.
- Never measured on a machine like yours, so no speed is shown.
Measured speed · decode, tokens per second
Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.
Architecture, for the curious
- model_type
- qwen2
- design
- softmax-GQA
- experts
- dense
- attention window
- none
- QK-norm
- no
- RoPE
- full
- norm
- RMSNorm, pre-norm
- activation
- SwiGLU
- tied head
- no
- modality
- text
- GPU-resident
- eligible