Structured output · 01 of 18 · measured 2026-09-02

A Go struct the model can't break

goinfer turns a Go struct's json tags into a grammar and holds the model to it while it writes. The output is always the struct's shape. What each field says is still the model's.

outputalways the struct's shape, so json.Unmarshal parses itnot guaranteedthat a value is true, or that a number fits its Go typecost of the mask1.21x a 6.2 ms decode step, on one PC, for one documenta parse rate on a real modelnot measured

The problem

Ask a language model for JSON and you usually get JSON. Usually. Sometimes there is a trailing comma, a key spelled differently, a number where you wanted a string, or a sentence of chat before the opening brace. Go's json.Unmarshal rejects each of these. The common answer is a retry loop: parse, fail, ask again, and hope. That costs time, and it never becomes a guarantee.

What goinfer does

You give the constrain package a Go value of your struct type. It reads the type and its json tags, derives a JSON Schema, and compiles the schema into a grammar. A model writes one token at a time (a token is a word or a piece of one). While it writes, goinfer removes every next token that would take the output outside that grammar. This is the example from the project's README:

type Person struct {
    Name string   `json:"name"`
    Age  int      `json:"age"`
    Tags []string `json:"tags"`
}

g, _ := constrain.GrammarFromStruct(Person{})       // struct → JSON Schema → grammar
sp.LogitProcessor = constrain.NewMasker(g, toks, eos).StopWhenComplete().Process

out := generate(sp)                                  // constrained decode
var p Person
_ = json.Unmarshal(out, &p)                          // shape guaranteed, not magnitude

generate, toks (each token's bytes) and eos (the token ids that end generation) are stand-ins there. The real calls are in examples/confidence. It uses a hand-written schema instead of a struct, and it also turns on per-field confidence, which How sure was it? covers. This is its generation half, unchanged:

g, err := constrain.JSONSchema([]byte(schema))
check(err, "schema")
mask := constrain.NewMasker(g, constrain.TokenBytes(m.Config().VocabSize, tok.TokenText), stop).
	StopWhenComplete().CaptureConfidence(constrain.ConfidenceOptions{})

ch, gen := m.Generate(context.Background(), ids, 128, decoder.SamplingParams{LogitProcessor: mask.Process, StopIDs: stop})
var out []int
for id := range ch {
	out = append(out, id)
}
check(gen.Err(), "generate")

To see what a struct becomes, constrain.SchemaFromStruct returns the schema itself. The package's tests check this exact output (ExampleSchemaFromStruct), for a Person with Name, Age and an Email string tagged json:"email,omitempty":

{"additionalProperties":false,"properties":{"age":{"type":"integer"},"email":{"type":"string"},"name":{"type":"string"}},"required":["name","age"],"type":"object"}

Fields without omitempty, and not pointers, are required. omitempty and pointer fields are optional.

How it works

At each step the model gives every token in its vocabulary a score. The Qwen2.5 models have 151,936 tokens. Masker.Process runs before the sampler, the step that picks the next token from those scores. For each token it asks the grammar whether that token's bytes keep the output a valid start of a document that fits the schema. If not, the token's score becomes minus infinity, so the sampler cannot pick it. Nothing is retried; the bad token is unreachable.

The grammar works on bytes, so a token that spans structure, like a closing quote and brace together, is judged correctly. Keys are matched against the properties not yet written, each at most once, in any order. The end-of-sequence token stays blocked until the document is complete, and StopWhenComplete then ends generation at the first complete document.

A schema keyword the compiler cannot enforce is an error when you build the grammar, so you never believe a constraint is in force when it is not.

What was measured

No record in the repo measures a parse rate on a real model. Nothing was counted as "N answers, N parsed". What does exist is tests of the grammar itself. They drive the masker with random choices over a small made-up vocabulary, not a model, and assert that it never dead-ends and that the output is valid:

whatwhat it assertswhere
5 schemas, 200 random generations eachoutput parses and conforms, checked by an independent validatorTestSchema_propertyValidates
a struct with nested and optional fields, 200 generationsoutput unmarshals into the struct with unknown fields disallowedTestGrammarFromStruct_roundTrip
3000 random walks over the plain JSON grammarevery finished document is valid per encoding/jsonTestConstrainedDecode_alwaysValidJSON

The cost of the mask was measured once, on 2026-09-02. It is recorded in the project's work queue, docs/QUEUE.md, in the entry on the cost of constrained decoding. The mask work is the same size whatever the model is. Timed at every step of a 17-token JSON document, with a 151,936-token vocabulary, it averaged 1.299 ms/step.

The record compares that with a decode step, the time to produce one token, for a 1.5B model (Qwen2.5-Coder-1.5B, 8-bit weights) held on the GPU of the Linux PC (Ryzen 7 3700X, RTX 2070 SUPER 8 GB). The record does not say which machine timed the mask itself. That step took 6.2 ms at the 64th token of context and 7.4 ms at the 512th. With the mask, constrained decoding took 1.21× the unconstrained time at 6.2 ms per step, and 1.18× at 7.4 ms. The slowest single grammar state would give at most 1.72×. That is an upper bound, not a typical step.

A large part of that came from one fix. Inside a JSON string, 96.88% of the vocabulary's tokens hold no quote, backslash or control byte. Those tokens are always allowed there and leave the grammar where it was, so goinfer answers them with one bit test instead of checking each byte against the grammar.

The record also says where this stops holding. One PC, one vocabulary and one small struct were timed. The mask does not get cheaper for a faster model, so against a step of about 2 ms it would be about 1.65×. And it left two costs unmeasured, both with one cause: a request that runs a mask like this one skips some faster decode paths. On setups that use those paths, the real cost is higher than these ratios.

Use it

go get github.com/townsendmerino/goinfer/decoder@latest github.com/townsendmerino/goinfer/tokenizer@latest

Then import github.com/townsendmerino/goinfer/constrain beside decoder, tokenizer and chat. The complete program to copy is examples/confidence. From a clone of the repo:

go run ./examples/confidence ~/models/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf "I was charged twice for order #123. - Ann"

The chat demo takes a schema file with --schema (from the README: go run ./demo/chat --model … --schema person.schema.json). The HTTP server, goinfer-serve, takes the same schema subset in a request's response_format, as {"type":"json_schema", …}. That path has no Go struct in it, so it takes a schema. Sketch, not run: a Go client can call constrain.SchemaFromStruct(Person{}) and send those bytes as the schema.

What it doesn't do

It doesn't check what the values say.
The grammar knows a field is a string or an integer, not whether it is correct. An age of 412 fits an int. Free text inside a string is the model's own. A special control token that is not a stop token can also appear there as plain text: a known audit finding, deferred and not fixed.
It doesn't bound a number's size.
JSON Schema integers have no width, so a uint8 field can be given 99999 and json.Unmarshal will return an error for it. Unsigned fields are stopped from going negative, and that is all.
It doesn't take every schema.
Supported: objects, arrays with minItems and maxItems, string, number, integer, boolean, null, enum, const, and minimum only when it is 0. Any other keyword (pattern, maxLength, oneOf, $ref, uniqueItems) is a compile error, not a silent skip. So are a freeform object with no declared properties, more than 64 properties or enum entries, and nesting deeper than 64.
It doesn't take every Go type.
Maps, interfaces and recursive types are refused, as is a type with its own UnmarshalJSON. Only json tags are read, so a struct has no way to say 'one of these three strings'; for that, write the schema by hand. A pointer field is optional, but the model cannot write null into it.
It can't finish a document the token budget won't fit.
The grammar holds the model to valid prefixes; it does not make it hurry. If generation stops at the request's token limit (max_tokens) before the closing brace, you get a valid prefix that will not parse. The server reports that stop as finish_reason length.

From the repo: constrain · docs/QUEUE.md · docs/server.md · README.md · examples/confidence