Connect a tool · from docs/integrations/claude-code.md

Claude Code → goinfer

Claude Code speaks the Anthropic Messages API, so it points at /v1/messages with three env vars. Verified end to end 2026-09-02 on the numbers below.

goinfer-serve -model coder=~/models/qwen2.5-7b-instruct-q4_k_m.gguf \
  -quant int4 -backend cuda -ctx 16384

ANTHROPIC_BASE_URL=http://127.0.0.1:8080 \
ANTHROPIC_AUTH_TOKEN=goinfer \
ANTHROPIC_MODEL=coder claude

All three env vars are required. Non-loopback binds refuse to start without -api-key.

What passed, and the numbers

Model class: Qwen2.5-7B-Instruct q4 → int4, CUDA-resident. A model that can hold an agent loop is the requirement, not just a context window: it must emit a tool call and then answer from the result. A 1.5B re-calls the same tool forever, which looks like a server bug and is not one.

measured
tool loop (glob → read → answer) 3 turns, 2.83 s, correct answer, ends stop_reason: end_turn
streamed tool call event order message_start → ping → content_block_start → content_block_delta → content_block_stop → message_delta → message_stop (no [DONE])
usage on a streamed tool call present (this was M-26; it used to appear on the plain text stream only)
TTFT, realistic agent turn (cold) 8.85 s for a 2,293-token turn (25 tool schemas + a 3.5 KB system prompt) ≈ 259 tok/s prefill
TTFT, the same turn once warm 0.42–0.58 s (prefix reuse; see below)
TTFT, small turn (2 schemas) 0.89 s

Provenance: RTX 2070 SUPER, NVIDIA driver 595.91.07, -quant int4 -backend cuda, greedy, warm, /v1/messages non-streaming for the loop and streaming for the event/usage rows; turn size from this server's own /v1/messages/count_tokens.

What to expect, and what will bite

Turn 1 is cold; turns 2+ are not. A resident model now reuses the part of the prompt already in its GPU KV and prefills only what changed, which is what an agent loop mostly adds (one tool call plus one tool result). Same box, same 25-schema turn, identical answers:

turn without reuse with reuse
1 (cold) 8.86 s 8.80 s
2 9.01 s 0.58 s
3 9.13 s 0.42 s
whole loop 27.00 s 9.82 s

Note the shape as much as the ratio: without reuse the per-turn cost grows with the conversation; with it, the cost tracks what you actually added. GOINFER_NO_RESIDENT_REUSE=1 turns it off.

Reuse is skipped whenever the prompt diverges from what the cache holds — editing an earlier message, or a second conversation on the same server — and that turn cold-prefills. Prefix reuse is per-model and single-conversation: two interleaved conversations will each cold-prefill as they alternate.

Tool calls on a ChatML-template model (Qwen2.5 above) cannot be malformed and cannot name a tool you did not send, with any number of tools: tool_choice: "any" forces a call to one of them, and auto constrains a call from the moment the model starts one, while still letting it answer in prose (docs/server.md, "What tool_choice actually constrains"; docs/tool-call-coverage.md lists which families this holds for). Two items this page used to list as known-open are closed: N-18 (any with 2+ tools did not force a call — closed 2026-09-24) and M-20 (Gemma-4's tool rendering — fixed 2026-09-02, 0e7a5956). Gemma 4's own call syntax is parsed, not constrained, so Qwen2.5 remains the better choice for tool work.

cache_control and metadata are accepted and ignored. thinking is honoured for a model whose chat template has a thinking control (Qwen3, Qwen3.5, Gemma 4 — docs/server.md, "Reasoning models (thinking)"): with enabled or adaptive the reply carries a thinking content block (empty signature) before the text, and with no thinking field the reasoning is dropped and only text is returned. Smoke-tested with Claude Code 2.1.284 (2026-09-30) against the Qwen3.5-0.8B: it parses the thinking stream, accepts the empty signature, and in a tool loop replays the thinking blocks with every turn; serve renders them into the prompt the way the model's own template would (kept for the loop in progress, dropped before the last query) and answers 200. The test model could not use tools reliably, so the loop itself was poor — a model-quality limit, not the protocol. If you would rather not have thinking, send thinking: {"type": "disabled"} or start serve with -thinking off.

Retiring this page

Per docs/tasks/task-embed-and-harness-ux.md §3.5, a recipe is retired when serve check covers what it says. goinfer-serve check <url> already covers the model list, streamed chat with usage, structured output, stop sequences and count_tokens; the tool-loop and agent-turn-TTFT rows above are what it does not cover yet.

Which listed model actually tool-calls under a real agent

R11 (docs/measurements/cold-user-2026-09-06-nobara-pc.md): serve check's "tools, OpenAI" row passed against a server running a 1.5B model — and a real agent (opencode) driving that exact server then printed a fake JSON tool call as prose, twice, instead of a real one. The gap is schema SIZE: a minimal one-tool schema is not what a real agent sends. check now has a second row, "tools, harness-scale" (a dozen tools with nested parameters, the shape opencode's own "build" agent sends), specifically to predict this before an operator hits it.

Measured, not guessed — goinfer-chat models' tools: line records each registry checkpoint's result from that row (2026-09-07, nobara-pc):

  • qwen2.5-coder-0.5b — minimal schema: ok; harness-scale: skip, too small.
  • phi3-mini-4k, granite-4.0-h-tiny, gpt-oss-20b — not yet run against the harness-scale row; goinfer-chat models says so plainly rather than guessing.

The closest existing evidence for a model class that DOES hold up is this page's own table above, from a different measurement pass (2026-09-02, a 25-tool-schema agent loop, not this registry's checkpoints): Qwen2.5-7B-Instruct completed a real glob → read → answer tool loop; the same run's own note is blunt about the size that does not: "A 1.5B re-calls the same tool forever, which looks like a server bug and is not one." Until a registry entry at that class is run through serve check's harness-scale row specifically, treat 7B-and-up as the size to reach for behind a real agent, and the registry's small checkpoints (0.5B currently measured, ~3-4B untested) as demo-scale: they answer directly, and skip rather than hallucinate a tool call, under a schema shaped like a real one.