Models / Qwen2 / Qwen2.5generated from capability-matrix.json

Qwen2 / Qwen2.5

Qwen2 and Qwen2.5, including the Coder models most of goinfer's numbers are measured on.

chatcodesafetensors, GGUF

Qwen2.5-Coder 0.5B

Get it

fits on MacBook Pro, M1 Pro, 16 GBfits on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull qwen2.5-coder-0.5b
repo
Qwen/Qwen2.5-Coder-0.5B-Instruct-GGUF
file
qwen2.5-coder-0.5b-instruct-q4_k_m.gguf
quant
q4_k_m
size
0.49 GB
sha256
1d9614638d18024d…

Good for, and what it needs

Code completion and small edits; the smallest checkpoint here that writes usable Go.

~1 GB resident at --quant int4; runs on any laptop

Tool calling

minimal schema: ok; harness-scale (12 tools): skip — too small (measured 2026-09-07, nobara-pc)

Decisions · /v1/systemone

Label scoring, unmeasured.

What hasn't been shown

  • Too small for an agent's full tool list: given twelve tools, it answered instead of calling one.
  • Never measured on a machine like yours, so no speed is shown.

How sure we are

Against the released model
100.0% picks the same next token as HuggingFace
1.00000 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

full-oracle 100.0%/1.00000 · what parity-gated means

Measured speed · decode, tokens per second

your pick · MacBook Pro, M1 Pro, 16 GBMetal · 2026-09-25
goinfer
171.2 Ollama
144.5
1.18× Ollama From the 2026-09-25 pre-registered sweep. The Metal decode kernels changed on 2026-09-26.
your pick · RTX 2070 SUPER, 8 GBCUDA · 2026-09-29
goinfer
367.9 Ollama
263.8
1.40× Ollama
your pick · Ryzen 7 3700X, CPU onlyCPU · 2026-09-29
goinfer
52 Ollama
57.4
0.91× Ollama

Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.

Qwen2.5-Coder 1.5B

Get it

fits on MacBook Pro, M1 Pro, 16 GBfits on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:q4_k_m
repo
Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
file
qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
quant
q4_k_m
size
1.1 GB
sha256
cc324af070c2ecbf…

Good for, and what it needs

Pinned as a release tier (pull/curated.json), not a registry checkpoint, so it carries no good-for, needs or tools fields.

Tool calling

not yet measured

Decisions · /v1/systemone

Label scoring, chat-v1, calibrated: top-1 0.2774, ECE 0.2102, over 872 out-of-distribution questions from a public decisions corpus. Measured 2026-09-28 on the RTX 2070 SUPER.

Label scoring reads the model's own probabilities for the option labels; it is not a trained decision head, and the record explains how far behind one it sits: the measurement.

What hasn't been shown

  • Tool calling hasn't been measured on Qwen2.5-Coder 1.5B.
  • Never measured on a machine like yours, so no speed is shown.

Measured speed · decode, tokens per second

your pick · MacBook Pro, M1 Pro, 16 GBMetal · 2026-09-26
goinfer
89.3 Ollama
84.3
1.06× Ollama After the R18b kernel change, a same-session A/B against Ollama. Not the pre-registered sweep.
your pick · RTX 2070 SUPER, 8 GBCUDA · 2026-09-29
goinfer
270.7 Ollama
183.7
1.47× Ollama
your pick · Ryzen 7 3700X, CPU onlyCPU · 2026-09-29
goinfer
24.3 Ollama
24
level with Ollama Every pair within 3% of Ollama (1.007 to 1.015).

Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.

Architecture, for the curious
model_type
qwen2
design
softmax-GQA
experts
dense
attention window
none
QK-norm
no
RoPE
full
norm
RMSNorm, pre-norm
activation
SwiGLU
tied head
no
modality
text
GPU-resident
eligible
registry description: Alibaba Qwen2/2.5 dense (q/k/v bias)