Structured output · 04 of 18 · measured 2026-09-27

How sure was it?

Ask for a JSON answer, and each choice, yes/no and count in it comes back with the model's own probability over the answers your schema allowed.

turn it on"goinfer_confidence": truecosts1–4% of each token's timefieldsenum, boolean, integercalibratedno

The problem

A constrained answer always looks sure of itself.

goinfer can already force a model's output into a JSON schema, so the reply always parses and every field holds a value the schema allows. That's useful, and it hides something. A model that knew the answer and a model that picked between two options by a hair produce the same clean JSON. Nothing in the reply tells them apart.

If you're sorting support tickets, or reading invoices, or deciding which requests an agent may act on alone, that difference is the one you care about.

What goinfer does

A model writes its answer one token at a time; a token is a word or a piece of one. At every step of a constrained answer, the model scores every possible next token, goinfer works out which of them the schema allows, and it throws the other scores away. With confidence turned on, it also reads what the model thought at the one point where each field was decided, and reports it next to the answer.

example · Qwen2.5-Coder 1.5B · triage schema
ticket
    Illustrative values, in the shape the server returns. The measured results are further down.

    That threshold is the intended use. The number ranks answers, and a low one is wrong more often, so you can route on it: handle the confident ones, and send the rest to a person.

    Why only the deciding token

    Most tokens in a constrained answer aren't a choice at all. Once a model has started writing "sh, the schema allows only shipping, so the rest of the word and its closing quote have probability 1.0. Averaging those in would make every enum look about two-thirds certain before the model had a say.

    "category":structure
    "shdecides
    ippingforced
    "forced

    An enum value is three tokens and one decision. goinfer reads the model at the free token and adds up the probability by the option each allowed token would spell. Where the tokens split is illustrative.

    So a field's confidence comes only from tokens where the schema left the model a real choice. A field the schema forces outright, like an enum with one option, isn't reported, because its value says nothing about the model. An integer reports the least certain of its digits.

    What was measured

    Before anything was built, the question was whether this number means anything. We wrote 60 support tickets whose correct labels follow from rules in the prompt. On 2026-09-27 we ran two models over them on a MacBook Pro (M1 Pro, 16 GB), on its GPU through Metal. The models were Qwen2.5-Coder 1.5B Instruct and Qwen2.5 7B Instruct, with weights stored in about 4 bits each (Q4_K_M quantization). Each always took its most likely next token (greedy decoding). The bar was set in writing before the graded runs (pre-registered): for each field kind, the confidence must separate right answers from wrong ones with an AUROC of at least 0.65. Below 0.55 the kind would fail; in between, it would be left undecided.

    AUROC, in plain terms: pick one right answer and one wrong answer at random. How often did the right one carry the higher confidence? 0.5 is a coin toss. 1.0 is perfect.

    1.5B: counts, enough mistakes to grade7B: too few mistakes to countshaded: below 0.55 would have failed
    field kind1.5B AUROC (right / wrong)7B AUROC (right / wrong)result
    enum0.847 (48 / 12)0.967 (55 / 5)reported
    boolean0.727 (38 / 22)0.905 (58 / 2)reported
    integer0.680 (43 / 17)0.942 (57 / 3)reported
    number— (60 / 0)— (60 / 0)not reported yet
    string— (57 / 3)— (60 / 0)not reported yet

    A model counted for a field kind only if it had at least 8 right and 8 wrong answers of that kind. The 7B got almost everything right, which is good for the 7B and useless for grading: a confidence can't be judged against mistakes that didn't happen. Its numbers point the same way and decide nothing. Integer cleared the bar on one model by 0.03, on 17 wrong answers. Numbers and free text need a harder test set, so they aren't reported until one exists.

    4.20%
    of a token's time on the 1.5B (0.570 ms of 13.56 ms). Median 1.88%.
    1.44%
    of a token's time on the 7B (0.542 ms of 37.55 ms). Median 0.72%.

    Those two figures are from the test, which read the model at every position. The version that shipped reads only the positions that decide a value and skips free text, where nearly every token is allowed. Its own cost test added 0.29 ms at an object key, 0.28 ms at an enum value and nothing inside a free string, about 2% (1.5B) and 0.8% (7B) of a token at the positions it reads. That test ran on a busy machine, so read it as indicative. It's off unless you ask for it, and the answer itself doesn't change: the same request without the flag returns the same content.

    Use it

    # with goinfer-serve -model <model.gguf> running, any /v1/chat/completions
    # or /v1/completions request with a json_schema
    curl -s localhost:8080/v1/chat/completions -d '{
      "messages": [{"role": "user", "content": "Triage this ticket: …"}],
      "response_format": {"type": "json_schema", "json_schema": {
        "name": "triage", "schema": { … }}},
      "goinfer_confidence": true
    }'
    
    # the response gains one record per reported field
    "goinfer_confidence": [
      {"path": "category", "kind": "enum", "value": "shipping",
       "confidence": 0.60, "free_tokens": 1, "calibrated": false,
       "distribution": {"billing": 0.02, "shipping": 0.60, "technical": 0.37,
                 "other": 0.01}},
      …
    ]

    If all you need is a pick between a few options, the server's decisions endpoint answers from a single read of the prompt, with no generation at all, in the request shape some decision APIs already use.

    What it doesn't do

    It isn't the probability that the answer is right.
    It's the model's probability over what the schema allowed, which is a narrower thing. It was tested for ranking (is a low number wrong more often?), not for calibration, and every field it reports says "calibrated": false. A 0.9 doesn't mean nine times in ten.
    It doesn't cover numbers or free text.
    Only enum, boolean and integer fields are reported. Number and string fields are not: the test set produced too few wrong numbers and names to judge them.
    It rests on one small test.
    60 support tickets, one kind of task, written by us, and two models from one family (Qwen2.5). A harder labelled set would say more, and might move the integer result either way.
    It doesn't work everywhere a schema does.
    The server refuses the flag with a 400, rather than dropping it silently, on requests with tools or images, on /v1/jobs and batches, and on /v1/responses and /v1/messages. It also turns off grammar-fused speculative decoding for that request: the speed-up that guesses several tokens ahead and checks them in one step.
    It doesn't make a small model good.
    A 1.5B-parameter model that is unsure is often right to be. The number tells you where to look, not what the answer should have been.

    From the repo: docs/server.md · docs/measurements/confidence-c0-2026-09-27.md · docs/tasks/parked/task-constrained-confidence.md · constrain/confidence.go · examples/confidence