Connect a tool · from docs/integrations/claude-code.md
Claude Code → goinfer
Claude Code speaks the Anthropic Messages API, so it points at /v1/messages with three env
vars. Verified end to end 2026-09-02 on the numbers below.
goinfer-serve -model coder=~/models/qwen2.5-7b-instruct-q4_k_m.gguf \
-quant int4 -backend cuda -ctx 16384
ANTHROPIC_BASE_URL=http://127.0.0.1:8080 \
ANTHROPIC_AUTH_TOKEN=goinfer \
ANTHROPIC_MODEL=coder claude
All three env vars are required. Non-loopback binds refuse to start without -api-key.
What passed, and the numbers
Model class: Qwen2.5-7B-Instruct q4 → int4, CUDA-resident. A model that can hold an agent loop is the requirement, not just a context window: it must emit a tool call and then answer from the result. A 1.5B re-calls the same tool forever, which looks like a server bug and is not one.
| measured | |
|---|---|
tool loop (glob → read → answer) |
3 turns, 2.83 s, correct answer, ends stop_reason: end_turn |
| streamed tool call event order | message_start → ping → content_block_start → content_block_delta → content_block_stop → message_delta → message_stop (no [DONE]) |
usage on a streamed tool call |
present (this was M-26; it used to appear on the plain text stream only) |
| TTFT, realistic agent turn (cold) | 8.85 s for a 2,293-token turn (25 tool schemas + a 3.5 KB system prompt) ≈ 259 tok/s prefill |
| TTFT, the same turn once warm | 0.42–0.58 s (prefix reuse; see below) |
| TTFT, small turn (2 schemas) | 0.89 s |
Provenance: RTX 2070 SUPER, NVIDIA driver 595.91.07, -quant int4 -backend cuda, greedy,
warm, /v1/messages non-streaming for the loop and streaming for the event/usage rows;
turn size from this server's own /v1/messages/count_tokens.
What to expect, and what will bite
Turn 1 is cold; turns 2+ are not. A resident model now reuses the part of the prompt already in its GPU KV and prefills only what changed, which is what an agent loop mostly adds (one tool call plus one tool result). Same box, same 25-schema turn, identical answers:
| turn | without reuse | with reuse |
|---|---|---|
| 1 (cold) | 8.86 s | 8.80 s |
| 2 | 9.01 s | 0.58 s |
| 3 | 9.13 s | 0.42 s |
| whole loop | 27.00 s | 9.82 s |
Note the shape as much as the ratio: without reuse the per-turn cost grows with the
conversation; with it, the cost tracks what you actually added. GOINFER_NO_RESIDENT_REUSE=1
turns it off.
Reuse is skipped whenever the prompt diverges from what the cache holds — editing an earlier message, or a second conversation on the same server — and that turn cold-prefills. Prefix reuse is per-model and single-conversation: two interleaved conversations will each cold-prefill as they alternate.
Tool calls on a ChatML-template model (Qwen2.5 above) cannot be malformed and cannot name a tool
you did not send, with any number of tools: tool_choice: "any" forces a call to one of them, and
auto constrains a call from the moment the model starts one, while still letting it answer in prose
(docs/server.md, "What tool_choice actually constrains"; docs/tool-call-coverage.md lists which
families this holds for). Two items this page used to list as known-open are closed: N-18 (any with
2+ tools did not force a call — closed 2026-09-24) and M-20 (Gemma-4's tool rendering — fixed
2026-09-02, 0e7a5956). Gemma 4's own call syntax is parsed, not constrained, so Qwen2.5 remains
the better choice for tool work.
cache_control and metadata are accepted and ignored. thinking is honoured for a model whose chat template has a
thinking control (Qwen3, Qwen3.5, Gemma 4 — docs/server.md, "Reasoning models (thinking)"): with enabled or
adaptive the reply carries a thinking content block (empty signature) before the text, and with no thinking field
the reasoning is dropped and only text is returned. Smoke-tested with Claude Code 2.1.284 (2026-09-30) against the Qwen3.5-0.8B: it parses the thinking stream, accepts the
empty signature, and in a tool loop replays the thinking blocks with every turn; serve renders them into the prompt the way the
model's own template would (kept for the loop in progress, dropped before the last query) and answers 200. The test model could
not use tools reliably, so the loop itself was poor — a model-quality limit, not the protocol. If you would rather not have
thinking, send thinking: {"type": "disabled"} or start serve with -thinking off.
Retiring this page
Per docs/tasks/task-embed-and-harness-ux.md §3.5, a recipe is retired when serve check covers
what it says. goinfer-serve check <url> already covers the model list, streamed chat with
usage, structured output, stop sequences and count_tokens; the tool-loop and agent-turn-TTFT
rows above are what it does not cover yet.
Which listed model actually tool-calls under a real agent
R11 (docs/measurements/cold-user-2026-09-06-nobara-pc.md): serve check's "tools, OpenAI" row
passed against a server running a 1.5B model — and a real agent (opencode) driving that exact
server then printed a fake JSON tool call as prose, twice, instead of a real one. The gap is
schema SIZE: a minimal one-tool schema is not what a real agent sends. check now has a second
row, "tools, harness-scale" (a dozen tools with nested parameters, the shape opencode's own
"build" agent sends), specifically to predict this before an operator hits it.
Measured, not guessed — goinfer-chat models' tools: line records each registry
checkpoint's result from that row (2026-09-07, nobara-pc):
qwen2.5-coder-0.5b— minimal schema: ok; harness-scale: skip, too small.phi3-mini-4k,granite-4.0-h-tiny,gpt-oss-20b— not yet run against the harness-scale row;goinfer-chat modelssays so plainly rather than guessing.
The closest existing evidence for a model class that DOES hold up is this page's own table
above, from a different measurement pass (2026-09-02, a 25-tool-schema agent loop, not this
registry's checkpoints): Qwen2.5-7B-Instruct completed a real glob → read → answer tool
loop; the same run's own note is blunt about the size that does not: "A 1.5B re-calls the same
tool forever, which looks like a server bug and is not one." Until a registry entry at that
class is run through serve check's harness-scale row specifically, treat 7B-and-up as the
size to reach for behind a real agent, and the registry's small checkpoints (0.5B currently
measured, ~3-4B untested) as demo-scale: they answer directly, and skip rather than hallucinate
a tool call, under a schema shaped like a real one.