Control · 03 of 18 · measured 2026-09-12
Stop means stop
If a client loops, or you just want it to stop, one command stops the tokens: for one generation, or for the whole server. The controls live on a separate channel from the one the client talks to.
The problem
A program that drives a model server can go wrong in boring ways: a retry loop, or a model that keeps asking for tool calls. Sooner or later you want it to stop, while the misbehaving client is still running.
Killing the server works, but it unloads the model. It also cannot stop one generation and leave the rest alone.
What goinfer does
A generation is one reply being written, a token (a word or a piece of one) at a time. Every generation has an id, the one already in the response (chatcmpl-... for chat). You can list the running ones and cancel one by id, or halt everything without a restart. The model stays loaded, so after a resume the next request runs without reloading it.
Start the server with a control socket:
goinfer-serve -model local=~/models/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf \
-admin-socket /tmp/goinfer-admin.sock
With the socket set, the same binary is the control tool:
goinfer-serve status
goinfer-serve ls # id, model, start time, tokens so far
goinfer-serve cancel chatcmpl-... "stuck in a loop"
goinfer-serve halt "operator stop"
goinfer-serve resume
If the socket is not at the default path (/run/goinfer/admin.sock, or under ~/Library/Application Support/goinfer/ on macOS), pass -admin-socket <path> before any reason text. Go's flag parser stops at the first plain word, so a reason typed first hides the flag.
The routes are plain HTTP, so curl works too:
curl --unix-socket /tmp/goinfer-admin.sock http://localhost/admin/generations
curl --unix-socket /tmp/goinfer-admin.sock -X POST http://localhost/admin/generations/chatcmpl-.../cancel \
-H 'Content-Type: application/json' -d '{"reason":"stuck in a loop"}'
curl --unix-socket /tmp/goinfer-admin.sock -X POST http://localhost/admin/halt \
-H 'Content-Type: application/json' -d '{"reason":"operator stop"}'
Cancel replies {"id": ..., "found": true}; found is false if the generation had already finished. Halt replies {"halted": true, "reason": ..., "quiesced_in_ms": <n>}, and only after every cancelled generation has stopped. Both need a reason.
What the client sees: a streaming reply ends with finish_reason: "cancelled" and one more event, goinfer_cancelled, carrying the id and reason, so it cannot be taken for a normal stop. A non-streaming reply gets a 499 error (a status that means the request was cancelled). While halted, every inference route answers 503 {"error":"halted","reason":...}, and GET /health reports halted.
Two other ways to halt, with no HTTP client: -halt-file <path> (checked every 250 ms; touch halts, rm resumes), and SIGUSR1 to halt and SIGUSR2 to resume. Windows has no SIGUSR1; it uses the file or the socket.
How it works
Every generation registers its cancel function under its id, in a registry of running generations. A cancel cancels that generation's Go context. The token loop already checks the context on every token, so the generation stops at the next token. Nothing asks the model to stop.
A halt does three things in order: set a flag, cancel every registered generation, and wait (up to 30 seconds) until the registry is empty. The flag is checked before the cap on concurrent requests (-max-inflight), so a halt does not wait for a free slot. It is checked again when a queued request is finally let in. Without that second check, requests queued for the model's one decode worker (the loop that generates tokens) would have started after the halt. The design notes record this as a bug found while building it.
Control lives on its own socket for a plain reason: a client that can reach /v1 should not also be able to reach the stop. Set -admin-socket and /admin/* is not registered on the TCP listener at all. A request there gets a 404, not a 403, so it does not confirm the routes exist. The socket has no API key; its file permissions (mode 0600, owner only) are the check. Without a socket, --allow-admin puts the routes on the TCP listener behind the API key, and the server refuses to start without a key.
What was measured
One record, dated 2026-09-12: the halt-time record.
| Machine | A MacBook, Apple Silicon (arm64), CPU backend |
| Model | Qwen2.5-Coder-0.5B-Instruct: a 4-bit Q4_K_M file, converted at load to 8-bit weights and 8-bit activations (int8int8) |
| Build | goinfer 4b979d0 |
| Test | TestServe_haltUnderLoad: 32 concurrent streaming chats, -max-inflight 32 -max-queue 32, then POST /admin/halt |
| Time from halt to every cancelled generation stopped | ~10ms, with and without -race (Go's race detector) |
| Outcome | 1 of 32 cancelled mid-stream; 31 of 32 refused with 503 before starting |
Only one request streams at a time, because this build runs one generation per model. The other 31 were queued, and the halt refuses each one when its turn comes. Later builds can run several generations of one model at once (see Batching that doesn't change the answer); the halt time has not been measured with that on.
The ~10ms is coarse. The server checks for an empty registry every 10 ms, so a reading of 10 means "within about one check". The record calls it a single-machine sample, not a pass/fail bar.
Two tests check behaviour, not time. TestServe_cancelByID needs a real .gguf model file. It cancels a 4096-token stream by id and checks the cancelled finish, the goinfer_cancelled event, and that the id leaves the registry. It then checks that the next greedy request (greedy: always take the most likely next token, so a repeat gives the same text) matches one made before the cancel. That shows the cancel left the model's state clean. TestServe_adminSocket needs no model: halt, status and resume work over the socket without a key, and /admin/halt on TCP returns 404. The model-backed tests skip unless GOINFER_SERVE_MODEL names a model file. The halt-file and signal triggers have no test; the design notes say they were tried by hand.
Use it
-admin-socket <path>puts/admin/*on a mode-0600 Unix socket and off TCP. Off by default.--allow-adminputs/admin/*on the TCP listener. It needs--api-key, and is ignored when-admin-socketis set.-halt-file <path>halts while the file exists. If youresumewhile the file is still there, the next check (within 250 ms) halts again.-halt-exit-code Nexits the process with code N after a halt, so a restarter can be told not to undo it.- These shipped in v0.18.0 (CHANGELOG, dated 2026-09-13). The server guide, docs/server.md, does not describe these routes yet. The design notes and the Serving section of docs/ARCHITECTURE.md, under Control, do.
What it doesn't do
- It doesn't sandbox what a client does with text it already has.
- goinfer produces tokens and, when asked, tool calls. The client runs the tools. A halt stops the tokens; it cannot recall a call the client already received, and it cannot undo anything. The design notes for this feature say so in their first section.
- The socket is protected by file permissions, nothing more.
- Anything that runs as the same user can open it, agent included. On the TCP listener, --allow-admin uses the same API key as /v1, so a client with the key can halt too. Separating users is a deployment job, and the systemd and launchd service files that would do it are planned but not built.
- The halt time was measured once.
- One run on the CPU of a MacBook (Apple Silicon), with a 0.5B model and 32 requests. The test logs the time and asserts no bound on it. Nothing measures it on Metal or CUDA, or with a large model.
- The rest of the plan isn't built.
- The design notes still list these as open: a lease that halts the server when nobody renews it, token and time budgets, cancel by session, a broker on the agent side that checks the lease before each tool call, and a drill that pulls every stop under load and times it. The flags for a lease and budgets do not exist.
- Some SDKs hide the reason.
- The cancel reason travels in an extra event. The Anthropic Python SDK drops that event, and the OpenAI Python SDK raises an error on it if its opt-in strict response validation is on (checked with anthropic-python 1.5.0 and openai-python 3.8.0). The finish reason still reads cancelled.
From the repo: docs/tasks/task-halt-2026-09.md · docs/measurements/kill-switch-quiescence-2026-09-12.md · CHANGELOG.md · docs/ARCHITECTURE.md · internal/serveapp/halt.go · internal/serveapp/admin.go · internal/serveapp/admin_socket.go