Use it from Go · from docs/use-from-go.md
Use it from Go
goinfer is a Go module. You import it, and the model runs inside your own process: no cgo, no daemon, no Python, and it cross-compiles like any other Go program. This page is the path for that use. The HTTP server is covered separately in server.md.
Package documentation is on pkg.go.dev.
The packages this page uses are decoder (load and generate), tokenizer, chat (prompt
templates) and constrain (structured output).
Every code block below is copied from a program under
examples/. CI builds and vets
those programs, and a test fails if a block here stops matching its source, so what you read
here compiles.
Install
Requires Go 1.27 or newer. From inside your own module (go mod init … first):
go get github.com/townsendmerino/goinfer/decoder@latest github.com/townsendmerino/goinfer/tokenizer@latest
Name the packages you import, as above. A go get of the bare module path, with no package,
records the requirement but does not fetch enough to build against, and the next build fails
with missing go.sum entry for module providing package …. One command like the one above is
enough for every other package in the module (chat, constrain, …).
You also need a model file. Any GGUF of a supported family
works; goinfer-chat pull qwen2.5-coder-0.5b fetches a small one.
The smallest program
examples/embed/main.go loads a GGUF, renders a chat prompt,
generates and prints. decoder.Load takes the path to the model file itself:
// Load takes the PATH TO THE .gguf FILE (or a directory of safetensors) — not a directory
// containing one, which is the guess a cold reader makes from the parameter name.
m, err := decoder.Load(path, decoder.Options{})
check(err, "load")
defer m.Close()
The zero decoder.Options loads the weights at f32. Options{Quant: "int4"} is what the
command-line tools default to, and uses about a quarter of the memory.
The tokenizer comes from the same file. PreTokenizerDecline names a pre-tokenizer the
checkpoint declares and goinfer does not implement. Text can then split into different tokens
than the reference tokenizer would produce, and the example says so rather than staying quiet:
tok, err := tokenizer.LoadGGUF(path)
check(err, "tokenizer")
if d := tok.PreTokenizerDecline(); d != "" {
fmt.Fprintln(os.Stderr, "!! tokenizer:", d)
}
Why chat.Detect comes before encoding. Model.Generate is a raw completion primitive: it
continues whatever token sequence it is given, and it has no idea of chat turns. An instruct
model was trained on its own template (special tokens around each turn), and fed a bare
question it tends to repeat itself instead of answering. chat.Detect reads the template the
checkpoint ships and returns a renderer for it. When it does not recognize the template
(chat.ErrUnknownTemplate), the example falls back to the raw prompt. A rendered template
already carries the model's own beginning-of-sequence marker, so only the fallback asks the
tokenizer to add one:
// buildPrompt renders the checkpoint's own chat template when Detect recognizes it (ErrUnknownTemplate
// means fall back to the raw prompt) and encodes it. A rendered template already carries the family's
// own BOS marker (Template.Render's contract, chat/chat.go:112-113); addBOS=true would double it on
// Gemma/Llama-3/Mistral (M-39, audit-2026-09-10). Only the raw-text fallback needs the tokenizer to
// prepend one.
func buildPrompt(tok *tokenizer.Tokenizer, prompt string) ([]int, error) {
text := prompt
rendered := false
if tmpl, terr := chat.Detect(chat.Meta{ChatTemplate: tok.ChatTemplate(), HasToken: tok.Has}); terr == nil {
text = tmpl.Render("", []chat.Turn{{Role: "user", Content: prompt}})
rendered = true
}
return tok.Encode(text, !rendered)
}
Generate returns a channel of token ids, which closes when generation ends, and a
*decoder.Generation: once the channel has closed, its Err reports an error that stopped
generation. The smallest example ignores it;
the next one checks it.
out, _ := m.Generate(context.Background(), ids, 256, decoder.SamplingParams{})
for id := range out {
s, err := tok.Decode([]int{id})
check(err, "decode")
fmt.Print(s)
}
fmt.Println()
Output shaped like a Go struct
examples/structured/main.go derives a JSON Schema from a
struct's json tags and constrains generation to it. A field is required unless it is a pointer
or has omitempty:
type Person struct {
Name string `json:"name"`
Age int `json:"age"`
Langs []string `json:"langs"`
Email string `json:"email,omitempty"` // omitempty (or a pointer) makes a field optional
}
The constraint is a logit mask. At every step, Masker.Process sets every token that would
break the schema to −∞, so an invalid token cannot be sampled at all; nothing is retried.
StopWhenComplete ends generation at the first complete document. The stop ids are the model's
end-of-sequence and end-of-turn tokens, which the masker allows only once the document is
complete:
// The stop ids end generation; the masker allows them only once the document is complete.
var stop []int
for _, id := range []int{tok.Special().EOS, tok.Special().EndOfTurn} {
if id >= 0 {
stop = append(stop, id)
}
}
for _, s := range tmpl.Stops().Strings {
if id, ok := tok.TokenID(s); ok {
stop = append(stop, id)
}
}
g, err := constrain.GrammarFromStruct(Person{}) // struct -> JSON Schema -> grammar
check(err, "grammar")
mask := constrain.NewMasker(g, constrain.TokenBytes(m.Config().VocabSize, tok.TokenText), stop).StopWhenComplete()
ch, gen := m.Generate(context.Background(), ids, 128, decoder.SamplingParams{LogitProcessor: mask.Process, StopIDs: stop})
var out []int
for id := range ch {
out = append(out, id)
}
check(gen.Err(), "generate")
raw, err := tok.Decode(out)
check(err, "decode")
var p Person
check(json.Unmarshal([]byte(raw), &p), "unmarshal") // the shape is guaranteed; a value can still be wrong
fmt.Printf("%+v\n", p)
What is guaranteed is the shape: the keys, which are required, and each value's JSON type.
What is not is a number's magnitude (JSON Schema's integer has no width, so a uint8 field
can be handed 99999 and json.Unmarshal returns an error), and that the values are correct. The
model can still fill in the wrong name. GrammarFromStruct refuses a type it cannot describe:
a recursive type, a struct with no exported fields, and a type with its own UnmarshalJSON
(unless it decodes from a string through UnmarshalText, which maps to a string).
Output shaped like a JSON Schema
constrain.JSONSchema compiles a schema you already have.
examples/confidence/main.go uses this one:
const schema = `{"type":"object","additionalProperties":false,"required":["category","urgent","orders"],"properties":{
"category":{"enum":["billing","technical","shipping","other"]},
"urgent":{"type":"boolean"},
"orders":{"type":"integer"}}}`
g, err := constrain.JSONSchema([]byte(schema))
check(err, "schema")
The supported subset:
- objects, with required and optional properties and
additionalProperties: false; - arrays, with
items,minItemsandmaxItems; string,number,integer,booleanandnull;enumandconst;- nesting to any depth.
An object may have up to 64 properties and an enum up to 64 values. A keyword or shape outside
the subset is an error from JSONSchema, so a constraint that could not be enforced is never
silently dropped. A property-based test in constrain checks that every constrained generation
validates against its schema.
How sure was it: CaptureConfidence
The same example turns on confidence capture on the masker, then asks for it per field once generation is done:
mask := constrain.NewMasker(g, constrain.TokenBytes(m.Config().VocabSize, tok.TokenText), stop).
StopWhenComplete().CaptureConfidence(constrain.ConfidenceOptions{})
fields, err := mask.FieldConfidence(out)
check(err, "confidence")
for _, f := range fields {
fmt.Printf("%-10s %-12s confidence %.2f", f.Path, f.Value, f.Confidence)
if f.Distribution != nil {
fmt.Printf(" %v", f.Distribution)
}
fmt.Println()
}
For each enum, boolean and integer field, FieldConfidence reports the model's probability over
the values the schema allowed, at the position that decided the value. A low number is wrong
more often (measurement), which makes it useful for
routing, such as "ask a person below 0.7".
It is not the probability that the value is right, and it is not calibrated: 0.9 does
not mean right nine times in ten. server.md has the full description; the server
exposes the same numbers with "goinfer_confidence": true.
An image in, a struct out (GLM-OCR)
examples/invoice/main.go reads a scanned invoice into a Go struct. The struct is the one source
for both halves of the request: constrain.TemplateFromStruct(Invoice{}) is the prompt, because GLM-OCR is prompted for extraction
with a JSON template (an object of blank values) and not a JSON Schema, and constrain.GrammarFromStruct(Invoice{}) is the
guarantee. Type the amounts as numbers (float64, int): the model answers a numeric column with a bare number, and a
string-typed field would leave the grammar only whitespace to emit (see docs/server.md). The image reaches the decoder through
multimodal.QwenPreprocess, aikit's vision.LoadGlmOcrVisionEncoder and Model.GenerateQwenVL, with the image block spliced into the
prompt by multimodal.SpliceImageBlock. Needs the GLM-OCR checkpoint (zai-org/GLM-OCR, about 2.7 GB); about a minute and a half
on a CPU. Over HTTP the same request is an image_url part plus response_format (docs/server.md); in the terminal it is
goinfer-chat --model ~/models/glm-ocr --image invoice.png --schema invoice.schema.json.
What v1.0 will bind
goinfer is pre-1.0, so any of this can still change. Which surfaces the v1.0 tag will hold to
semver was decided on 2026-08-18 in api-tiers.md. Its Hard tier covers what this
page uses from decoder, tokenizer and chat:
decoder.Load,decoder.Options(withQuant),Model.Generate,CloseandConfig,decoder.SamplingParamsanddecoder.Generation;- the
tokenizer.Tokenizermethods used here (Encode,EncodeSegments,Decode,TokenText,TokenID,Has,Special,ChatTemplate) andtokenizer.LoadGGUF; chat.Detect,chat.Meta,chat.ErrUnknownTemplate,chat.Template'sRenderandRenderSegments, and thechat.Turnandchat.Stopstypes;- from
constrain:NewMasker, theMasker'sProcessandStopWhenComplete,TokenBytesandSchemaFromStruct.
Three things this page uses are in neither of that page's lists today: GrammarFromStruct,
JSONSchema, and the confidence API (CaptureConfidence, FieldConfidence). Until
api-tiers.md places them, treat them as not yet promised. Two smaller ones are unlisted too:
Tokenizer.PreTokenizerDecline and the Template.Stops method.
Run the examples
go run ./examples/embed ~/models/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf "Reverse a string in Go"
go run ./examples/structured ~/models/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf "Ann Lee, 34, writes Go and Rust"
go run ./examples/confidence ~/models/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf "I was charged twice for order #123. - Ann"
go run ./examples/invoice ~/models/glm-ocr testdata/glm_ocr/invoice.png # GLM-OCR checkpoint, not a GGUF