What's different

Things goinfer does that are unusual.

Each one has a writeup: the problem, what goinfer does, what was measured, and what it doesn't do. A writeup is published only after the owner has reviewed it.

Structured output

A Go struct the model can't break

Give goinfer a Go struct and the model can only write JSON that fits it, so json.Unmarshal parses it. It fixes the shape, not whether the values are right.

Limit It doesn't check what the values say.

Read →
Agents

Tool calls that can't come out malformed

When a model starts a tool call, goinfer holds it to valid JSON, a tool you supplied, and that tool's argument schema. It does not pick the tool for the model.

Limit It doesn't choose the right tool.

Read →
Control

Stop means stop

Cancel one generation by id, or halt them all from an HTTP call, a file or a signal. Control sits on its own Unix socket, away from the API.

Limit It doesn't sandbox what a client does with text it already has.

Read →
Structured output

How sure was it?

Each enum, boolean and integer field of a constrained JSON answer comes back with the model's own probability over the answers the schema allowed.

Limit It isn't the probability that the answer is right.

Read →
Structured output

Decisions without generating

goinfer answers a choice among a few options from one read of the prompt, with no generation. It ships; its accuracy on a real model is not yet measured.

Limit It hasn't been graded on a real model.

Read →
Correctness

Checked against the reference

Each model family is compared with HuggingFace's own implementation, the result is printed on the Models page, and the label says how strong the check was.

Limit It doesn't cover every quantization on every machine.

Read →
Memory

It refuses rather than swaps

If a model won't fit, goinfer says so and by how much before it loads, watches swap while it loads, and retries dense models with weight streaming.

Limit It doesn't hold the machine on its own.

Read →
Memory

A 26B model on an 8 GB card

Gemma 4 26B-A4B runs on an 8 GB card by streaming its experts from host RAM: 39.3 tokens/s on 2026-09-29. An architecture comparison, not like-for-like.

Limit It doesn't fit in a small amount of RAM.

Read →
Speed

Turn nine in under half a second

goinfer keeps several conversations' history ready, so a new turn reads only what was added. On one CUDA replay, turn 9 of an agent session began in 319.5 ms.

Limit It doesn't keep every conversation.

Read →
Serving

Work that survives a disconnect

A generation submitted as a job keeps running after the client leaves, and a new connection can replay it. With -job-dir, a restart records what happened.

Limit It doesn't keep the answer through a restart.

Read →
Distribution

One file, model inside

One file holds the runtime and a fixed Qwen2.5-Coder model, 0.5B or 1.5B. It is large, uses a GPU only on Linux with NVIDIA, and its speed figures are old.

Limit It isn't small.

Read →
Distribution

No toolchain of any kind

goinfer builds with CGO_ENABLED=0: no C compiler, CUDA toolkit or Python to build or run it, and CUDA and Metal are reached without cgo.

Limit It doesn't need nothing.

Read →
Concurrency

Batching that doesn't change the answer

Up to four conversations share each GPU step, and every reply is the one it would get alone. Measured on Metal, CUDA and CPU; a lone request is unchanged.

Limit It isn't continuous batching or paged attention.

Read →
Speed

Faster, with the same words

On Metal, guessing ahead from text already in context: 2.082× plain decode on copy-heavy requests, 1.068× on chat, and every temperature-0 reply unchanged.

Limit It doesn't help ordinary chat much.

Read →
Reproducibility

An upgrade that can't change your answers

One command builds two versions of goinfer, compares every score the model outputs byte for byte, and says IDENTICAL or shows the first difference.

Limit It doesn't hold across machines or operating systems.

Read →
Measurement

Numbers with their receipts

Every speed names its machine, checkpoint, quantization, versions, date and machine state, and both engines are timed over their own HTTP servers, in turn.

Limit It doesn't cover more than a few machines.

Read →
Loading

Starts without converting the model again

After a one-time conversion, a model loads by mapping a file: 0.003 s warm, about 1.3 s cold, against 9.16 s for a 7B. Generation speed is unchanged.

Limit It doesn't skip the first-load conversion.

Read →
Agents

Find out before your agent does

serve check sends a dozen-tool schema to a running model before you set up opencode, and fit sizes a checkpoint for your machine first. Both are measured.

Limit It doesn't guarantee a long agent session.

Read →