Serving · from docs/flags.md
serve flag reference
All 60 flags of serve, grouped, with the type and default of each. This page is a skim: one sentence per flag. serve --help carries the full text of every flag, including the trade-off each one makes and what was measured. Where a flag also has an environment variable, see the environment variable registry.
Every flag takes one or two dashes (-model and --model are the same). A flag marked also in goinfer-chat is shared with that binary.
Subcommands
Besides serving, the binary has subcommands, each with its own -h:
| Command | What it does |
|---|---|
serve pull <ref> |
Fetch a model, sha256-verified, into the local cache. |
serve check |
Drive a running server through a harness conversation and print per-feature verdicts. |
serve status | ls | cancel <id> | halt [reason] | resume |
Control a running server over its --admin-socket. |
serve --version |
Print the version and the backends compiled into this binary. |
Choosing models
What to serve. --model is the one flag most invocations need.
| Flag | Type | Default | What it does |
|---|---|---|---|
--model |
[name=]path-or-ref (repeatable) |
— | What to serve: a .gguf or .giw file, a model directory, or a reference fetched on first use (hf:owner/repo:quant, hf:owner/repo:safetensors for a safetensors checkpoint, demo:<tier>). Repeatable; name=path serves several models from one process, with optional per-model overrides after the path. More. |
--served-model-name |
string |
— | The id /v1/models reports for a single unnamed --model (default: the file or directory name). |
--lora |
string |
— | A PEFT LoRA adapter directory, merged into a safetensors base at load. Also in goinfer-chat. |
--adapter |
serveName=baseName=dir (repeatable) |
— | A compute-time LoRA adapter that shares a base model's resident weights. Repeatable. Needs a dense safetensors base; not compatible with --stream-weights. |
--vision |
string |
— | Vision tower directory for a multimodal --model (found automatically per family); enables image content parts. Defaults to the --model directory when it holds a tower. |
--vision-quant |
f32 | int8 |
f32 |
Vision encoder precision: f32 (bit-exact) or int8, which only speeds up the vision encoder on AVX512-VNNI CPUs. |
--vision-max-pixels |
int (pixels) |
0 |
GLM-OCR only: lower the image pixel budget to this many pixels. 0 (the default) keeps the model's own ceiling, 4.82 MP; it is never raised. The CPU vision tower's cost is superlinear in the image. |
--embed-model |
string |
— | An embedding model, served on /v1/embeddings: a directory (CodeRankEmbed / NomicBert layout), a decoder-as-embedder .gguf, or hf:<owner>/<repo>:safetensors (a NomicBert encoder, fetched and checked before its weights) or hf:<owner>/<repo>:<quant> (a GGUF). |
--embed-quant |
f32 | q8 |
f32 |
Embedding weight precision. |
--embed-served-model-name |
string |
— | The id /v1/models reports for the embedding model (default: the directory name). |
Other
Everything else.
| Flag | Type | Default | What it does |
|---|---|---|---|
--no-selftest |
bool |
false |
Skip the startup self-tests. They check each kernel the process will use against a reference on a small fixed input, and step a CPU kernel tier down (or decline a GPU backend) on a mismatch. On by default; check --hardware prints what they found. |
--version |
bool |
false |
Print the version, the backends compiled into this binary, and the Go toolchain, then exit. |
--decisions-template |
chat-v1 | bare-v1 |
chat-v1 |
Prompt template for POST /v1/systemone: chat-v1 (the model's chat template) or bare-v1 (no chat template). |
--decisions-calibration |
string |
— | A calibration.json of per-kind temperatures for /v1/systemone. Without it, answers are uncalibrated. |
--log-requests |
bool |
false |
Write one line to stderr per generation request when it finishes: route, model, status, prompt and completion tokens, time to first token, total time. Covers /v1/chat/completions, /v1/completions, /v1/responses and /v1/messages. Off by default; a request that failed before generating shows - for the model and token counts. |
Memory and swap safety
Read these before loading a model that is close to your RAM or VRAM.
| Flag | Type | Default | What it does |
|---|---|---|---|
--fit |
on | off |
true |
Size an unpinned load to what this machine actually has instead of a flat historical default, for example a larger CUDA context when the card has the free VRAM. --fit=off restores the fixed defaults. Also in goinfer-chat. |
--stream-weights |
bool |
false |
Run a model bigger than your RAM: page weights from disk instead of holding them resident. The flag to know before loading something large. Also in goinfer-chat. |
--weight-cache |
number |
0 |
Resident expert-weight budget in GB for --stream-weights. 0 = automatic, about half of available RAM. Also in goinfer-chat. |
--accept-slow |
bool |
false |
Load a --stream-weights paged-MoE model even when its predicted decode rate is below the 2 tok/s floor. Without this, such a load is refused and the predicted rate is named. Also in goinfer-chat. |
--moe-cache-experts |
bool |
false |
Run a MoE model whose experts exceed VRAM or RAM by streaming routed experts to the GPU per token. Bit-identical to fully resident, at the cost of a transfer per token. Off by default. Also in goinfer-chat. |
--moe-cache-slots |
int |
0 |
Per-layer expert slots to keep resident for a paged MoE model. On CUDA this is an upper bound: it is lowered if free VRAM cannot hold it. Also in goinfer-chat. |
--moe-pager |
mmap | pool |
pool on macOS, mmap elsewhere |
Backing mode for the expert pager of a .giw-paged MoE model: mmap (zero-copy, but on macOS the memory budget is not enforced) or pool (owned buffers, a firm cap on every platform). Also in goinfer-chat. |
--require-backend |
bool |
false |
Exit non-zero at startup if a model did not get the requested --backend's fast paths (no resident decode, or a prefill that fell back to the per-token loop). Under --backend auto it also refuses to start when auto passed over a GPU backend this binary has (no device answered), rather than run on the CPU; --backend cpu runs strict on the CPU. For batch jobs that should fail at second zero rather than run slow. |
Backend, precision and context
Which compute backend runs the model, at what precision, with how much context.
| Flag | Type | Default | What it does |
|---|---|---|---|
--backend |
auto | cpu | webgpu | cuda | metal |
auto |
The compute backend. auto (the default) picks cuda when a device answers, else metal on Apple silicon for int4 models (Metal re-quantizes other precisions, so those stay on the CPU), else cpu, and prints one line saying which. It never picks webgpu. --version lists the backends in this binary; naming one that is not there falls back to cpu. Also in goinfer-chat. |
--quant |
int4 | int4mix | q4k | int8int8 | int8 | f32 |
int4 |
Weight precision, the accuracy, speed and RAM trade-off. int4 is the default; f32 is unquantized. A prequantized .giw carries its own. More. Also in goinfer-chat. |
--embed-int4 |
bool |
true |
With --quant int4, keep the token-embedding and LM-head table at int4 as well, halving the largest resident tensor on a small model with a big vocabulary. On by default, except when the backend is metal (named, or chosen by auto), where the resident runner does not take an int4 table yet and the default is off. --embed-int4=false keeps it at int8. More. Also in goinfer-chat. |
--kv |
f32 | f16 | i8 |
f32 |
KV cache precision: f32 is bit-exact; f16 and i8 are lossy and smaller (f16 applies to GPU-resident models only). Also in goinfer-chat. |
--kv-quant |
f32 | i8 |
— | Deprecated: use --kv. Overrides the CPU KV cache precision alone. |
--ctx |
int |
0 |
GPU-resident KV capacity in positions. 0 keeps the backend default, which --fit may raise. Also in goinfer-chat. |
--direct-load |
bool |
off (on when GOINFER_GGUF_DIRECT is set) |
Load a plain .gguf straight into memory instead of through its sidecar .giw cache. Also enabled by setting GOINFER_GGUF_DIRECT. Also in goinfer-chat. |
--exact-prefill |
bool |
false |
Force bit-exact prompt ingestion on every backend, turning off the faster prefill paths that are not bit-identical. For diffing outputs across versions or reproducing a bug. Also in goinfer-chat. |
--cpu-exact-prefill |
bool |
false |
The same, for the CPU backend only: the f64-accumulating attention kernel instead of the faster f32 default. Also in goinfer-chat. |
Concurrency and queues
How many requests run at once, and what happens to the rest.
| Flag | Type | Default | What it does |
|---|---|---|---|
--max-concurrent |
int |
4 |
How many generations of one model run at once, capped by --kv-sessions. GPU-resident models run several only where the backend can batch their decode; vision and weight-streaming models, and --spec without --spec-adaptive, run one. |
--max-queue |
int |
8 |
Requests that may wait per model before new ones get a 429. 0 = unbounded. |
--max-inflight |
int |
128 |
Global cap on concurrent inference requests, applied before the per-model queue. A full cap returns 503 with Retry-After. 0 = unbounded. |
--cpu-batch |
auto | on | off |
— | Whether concurrent CPU generations of one model join their decode tokens into one batched forward (replies are bit-identical either way). auto batches models with at least 2 GiB of dense weights. |
--prefill-chunk |
int |
512 |
On a GPU-resident model serving several generations, prefill a long prompt that arrives mid-decode in chunks of this many tokens, with a decode step between chunks. 0 = off. |
--max-body-bytes |
int |
0 |
Cap on request-body size in bytes; a larger body is rejected with 413 before it is read. 0 derives the cap from the model's context window. |
Sessions and KV cache
Keeping conversations warm between requests, and across restarts.
| Flag | Type | Default | What it does |
|---|---|---|---|
--kv-sessions |
int |
4 |
Conversations kept prefilled in RAM for prompt-prefix reuse. 0 disables. On GPU backends this is also how many resident KV slots a model keeps, clamped by its memory guard. Metal keeps 2 slots unless this flag is given (each slot's KV is resident from the first token on unified memory). |
--session-dir |
string |
— | Directory that persists and restores KV sessions across restarts. |
--kv-idle-demote |
duration |
0s |
Demote a warm session's KV to --session-dir once it has been idle this long (for example 10m); it faults back in on the next matching request. 0 = off. |
--kv-demoted-max |
int |
64 |
Maximum demoted (on-disk) sessions to keep; older ones are dropped. Only with --kv-idle-demote. |
--job-dir |
string |
— | Directory for a durable job journal: one line per generation state change. After a restart, a job still running at its last recorded change is marked interrupted. Off by default. |
Speculative decoding
Lossless drafting: output is identical to plain decoding, only the speed changes.
| Flag | Type | Default | What it does |
|---|---|---|---|
--spec |
ngram |
— | Lossless n-gram (prompt-lookup) speculative decoding. Wins on copy-heavy traffic such as code edits, RAG and agent loops. Serves one generation at a time unless --spec-adaptive is also set. More. |
--spec-adaptive |
bool |
false |
Experimental. With --spec ngram, keeps concurrent generations: one speculates only while it is alone and joins batched decode when others arrive. Needs a resident whose decode can batch. Graded on CUDA: faster than batching alone, but under load slower than batching on chat (0.86x) and slower than plain --spec ngram on copy (0.80x), so not recommended there (docs/measurements/mc4-candidate-cuda-2026-10-01.md). |
--drafter |
string |
— | A pretrained block-drafter directory (DFlash) paired with --model: it proposes a block of tokens per round and the target verifies them in one pass. Lossless; greedy requests only; needs a resident GPU backend. More. |
Network and security
Where the server listens and who may talk to it.
| Flag | Type | Default | What it does |
|---|---|---|---|
--addr |
string |
127.0.0.1:8080 |
Listen address. Loopback by default; bind 0.0.0.0 only with --api-key and TLS, or behind a TLS-terminating proxy. |
--api-key |
string |
— | Shared secret every request must send as Authorization: Bearer <key> or x-api-key. Falls back to $GOINFER_API_KEY. Required to bind anywhere but loopback, and with --allow-admin. |
--tls-cert |
string |
— | PEM certificate file. With --tls-key, serves HTTPS instead of plaintext HTTP. |
--tls-key |
string |
— | PEM private key file, paired with --tls-cert. |
--web |
bool |
false |
Serve a local browser UI at / to chat, and to pull and load GGUF files or whole safetensors checkpoints from Hugging Face. Off by default: its pull route starts multi-gigabyte downloads. |
--allow-admin |
bool |
false |
Enable /admin/* on the TCP listener: model load and unload, cancel, halt, resume. A deliberate opt-in that requires --api-key. Ignored when --admin-socket is set. More. |
Admin and halt
Controlling a running server without restarting it.
| Flag | Type | Default | What it does |
|---|---|---|---|
--admin-socket |
string |
— | Serve /admin/* on a Unix socket (mode 0600, no key check: file permissions are the auth) instead of TCP. The same binary controls it with status, ls, cancel, halt and resume. |
--halt-file |
string |
— | Poll this path every 250 ms: present halts the server (inference routes return 503 and in-flight generations are cancelled), absent resumes it. A supervisor halts with touch and resumes with rm. |
--halt-exit-code |
int |
0 |
When nonzero, a halt exits the process with this code once every cancelled generation has stopped, for supervisors whose restart policy must not undo a deliberate halt. 0 = a halt never exits. |
--unload-drain-wait |
duration |
5s |
How long POST /admin/models/unload waits for in-flight requests to drain before returning 202. |
Reasoning models
Models that think before they answer (Qwen3, Qwen3.5, Gemma 4): whether they are prompted to, and where the thinking goes in the reply.
| Flag | Type | Default | What it does |
|---|---|---|---|
--thinking |
template | asis | on | off |
template |
Default thinking mode for a model whose chat template has a recognised thinking control (Qwen3, Qwen3.5, Gemma 4): template (the default: what the checkpoint's own template renders), asis (the prompt serve rendered before thinking was modelled: nothing written, the model decides), on, or off. A request overrides it. |
--reasoning-format |
deepseek | deepseek-legacy | none |
deepseek |
How a reply's reasoning reaches the client: deepseek (clean content, reasoning in reasoning_content / Anthropic thinking blocks), deepseek-legacy (both, tags kept in content), or none (nothing separated). |
--reasoning-budget |
auto | unlimited | N |
auto |
A ceiling on how long a thinking reply may think before serve forces the block closed, so the reply always has room to answer: auto (the default: at most three quarters of the request's max_tokens), unlimited, or N tokens. A request's thinking_token_budget / Anthropic budget_tokens applies too. Off under -spec / -drafter. |
--tool-format |
auto | hermes | template |
auto |
How tools are put in the prompt for a model whose own template declares a call form goinfer can render exactly (Qwen3.5's XML, Gemma 4's canonical template): auto (the default: each family's measured default — the model's own form for Gemma 4, goinfer's prompt for Qwen3.5), hermes (goinfer's own prompt) or template (the model's own; a tool_choice naming a function is then a 400). |