Serving · from docs/flags.md

serve flag reference

All 60 flags of serve, grouped, with the type and default of each. This page is a skim: one sentence per flag. serve --help carries the full text of every flag, including the trade-off each one makes and what was measured. Where a flag also has an environment variable, see the environment variable registry.

Every flag takes one or two dashes (-model and --model are the same). A flag marked also in goinfer-chat is shared with that binary.

Subcommands

Besides serving, the binary has subcommands, each with its own -h:

Command What it does
serve pull <ref> Fetch a model, sha256-verified, into the local cache.
serve check Drive a running server through a harness conversation and print per-feature verdicts.
serve status | ls | cancel <id> | halt [reason] | resume Control a running server over its --admin-socket.
serve --version Print the version and the backends compiled into this binary.

Choosing models

What to serve. --model is the one flag most invocations need.

Flag Type Default What it does
--model [name=]path-or-ref (repeatable) — What to serve: a .gguf or .giw file, a model directory, or a reference fetched on first use (hf:owner/repo:quant, hf:owner/repo:safetensors for a safetensors checkpoint, demo:<tier>). Repeatable; name=path serves several models from one process, with optional per-model overrides after the path. More.
--served-model-name string — The id /v1/models reports for a single unnamed --model (default: the file or directory name).
--lora string — A PEFT LoRA adapter directory, merged into a safetensors base at load. Also in goinfer-chat.
--adapter serveName=baseName=dir (repeatable) — A compute-time LoRA adapter that shares a base model's resident weights. Repeatable. Needs a dense safetensors base; not compatible with --stream-weights.
--vision string — Vision tower directory for a multimodal --model (found automatically per family); enables image content parts. Defaults to the --model directory when it holds a tower.
--vision-quant f32 | int8 f32 Vision encoder precision: f32 (bit-exact) or int8, which only speeds up the vision encoder on AVX512-VNNI CPUs.
--vision-max-pixels int (pixels) 0 GLM-OCR only: lower the image pixel budget to this many pixels. 0 (the default) keeps the model's own ceiling, 4.82 MP; it is never raised. The CPU vision tower's cost is superlinear in the image.
--embed-model string — An embedding model, served on /v1/embeddings: a directory (CodeRankEmbed / NomicBert layout), a decoder-as-embedder .gguf, or hf:<owner>/<repo>:safetensors (a NomicBert encoder, fetched and checked before its weights) or hf:<owner>/<repo>:<quant> (a GGUF).
--embed-quant f32 | q8 f32 Embedding weight precision.
--embed-served-model-name string — The id /v1/models reports for the embedding model (default: the directory name).

Other

Everything else.

Flag Type Default What it does
--no-selftest bool false Skip the startup self-tests. They check each kernel the process will use against a reference on a small fixed input, and step a CPU kernel tier down (or decline a GPU backend) on a mismatch. On by default; check --hardware prints what they found.
--version bool false Print the version, the backends compiled into this binary, and the Go toolchain, then exit.
--decisions-template chat-v1 | bare-v1 chat-v1 Prompt template for POST /v1/systemone: chat-v1 (the model's chat template) or bare-v1 (no chat template).
--decisions-calibration string — A calibration.json of per-kind temperatures for /v1/systemone. Without it, answers are uncalibrated.
--log-requests bool false Write one line to stderr per generation request when it finishes: route, model, status, prompt and completion tokens, time to first token, total time. Covers /v1/chat/completions, /v1/completions, /v1/responses and /v1/messages. Off by default; a request that failed before generating shows - for the model and token counts.

Memory and swap safety

Read these before loading a model that is close to your RAM or VRAM.

Flag Type Default What it does
--fit on | off true Size an unpinned load to what this machine actually has instead of a flat historical default, for example a larger CUDA context when the card has the free VRAM. --fit=off restores the fixed defaults. Also in goinfer-chat.
--stream-weights bool false Run a model bigger than your RAM: page weights from disk instead of holding them resident. The flag to know before loading something large. Also in goinfer-chat.
--weight-cache number 0 Resident expert-weight budget in GB for --stream-weights. 0 = automatic, about half of available RAM. Also in goinfer-chat.
--accept-slow bool false Load a --stream-weights paged-MoE model even when its predicted decode rate is below the 2 tok/s floor. Without this, such a load is refused and the predicted rate is named. Also in goinfer-chat.
--moe-cache-experts bool false Run a MoE model whose experts exceed VRAM or RAM by streaming routed experts to the GPU per token. Bit-identical to fully resident, at the cost of a transfer per token. Off by default. Also in goinfer-chat.
--moe-cache-slots int 0 Per-layer expert slots to keep resident for a paged MoE model. On CUDA this is an upper bound: it is lowered if free VRAM cannot hold it. Also in goinfer-chat.
--moe-pager mmap | pool pool on macOS, mmap elsewhere Backing mode for the expert pager of a .giw-paged MoE model: mmap (zero-copy, but on macOS the memory budget is not enforced) or pool (owned buffers, a firm cap on every platform). Also in goinfer-chat.
--require-backend bool false Exit non-zero at startup if a model did not get the requested --backend's fast paths (no resident decode, or a prefill that fell back to the per-token loop). Under --backend auto it also refuses to start when auto passed over a GPU backend this binary has (no device answered), rather than run on the CPU; --backend cpu runs strict on the CPU. For batch jobs that should fail at second zero rather than run slow.

Backend, precision and context

Which compute backend runs the model, at what precision, with how much context.

Flag Type Default What it does
--backend auto | cpu | webgpu | cuda | metal auto The compute backend. auto (the default) picks cuda when a device answers, else metal on Apple silicon for int4 models (Metal re-quantizes other precisions, so those stay on the CPU), else cpu, and prints one line saying which. It never picks webgpu. --version lists the backends in this binary; naming one that is not there falls back to cpu. Also in goinfer-chat.
--quant int4 | int4mix | q4k | int8int8 | int8 | f32 int4 Weight precision, the accuracy, speed and RAM trade-off. int4 is the default; f32 is unquantized. A prequantized .giw carries its own. More. Also in goinfer-chat.
--embed-int4 bool true With --quant int4, keep the token-embedding and LM-head table at int4 as well, halving the largest resident tensor on a small model with a big vocabulary. On by default, except when the backend is metal (named, or chosen by auto), where the resident runner does not take an int4 table yet and the default is off. --embed-int4=false keeps it at int8. More. Also in goinfer-chat.
--kv f32 | f16 | i8 f32 KV cache precision: f32 is bit-exact; f16 and i8 are lossy and smaller (f16 applies to GPU-resident models only). Also in goinfer-chat.
--kv-quant f32 | i8 — Deprecated: use --kv. Overrides the CPU KV cache precision alone.
--ctx int 0 GPU-resident KV capacity in positions. 0 keeps the backend default, which --fit may raise. Also in goinfer-chat.
--direct-load bool off (on when GOINFER_GGUF_DIRECT is set) Load a plain .gguf straight into memory instead of through its sidecar .giw cache. Also enabled by setting GOINFER_GGUF_DIRECT. Also in goinfer-chat.
--exact-prefill bool false Force bit-exact prompt ingestion on every backend, turning off the faster prefill paths that are not bit-identical. For diffing outputs across versions or reproducing a bug. Also in goinfer-chat.
--cpu-exact-prefill bool false The same, for the CPU backend only: the f64-accumulating attention kernel instead of the faster f32 default. Also in goinfer-chat.

Concurrency and queues

How many requests run at once, and what happens to the rest.

Flag Type Default What it does
--max-concurrent int 4 How many generations of one model run at once, capped by --kv-sessions. GPU-resident models run several only where the backend can batch their decode; vision and weight-streaming models, and --spec without --spec-adaptive, run one.
--max-queue int 8 Requests that may wait per model before new ones get a 429. 0 = unbounded.
--max-inflight int 128 Global cap on concurrent inference requests, applied before the per-model queue. A full cap returns 503 with Retry-After. 0 = unbounded.
--cpu-batch auto | on | off — Whether concurrent CPU generations of one model join their decode tokens into one batched forward (replies are bit-identical either way). auto batches models with at least 2 GiB of dense weights.
--prefill-chunk int 512 On a GPU-resident model serving several generations, prefill a long prompt that arrives mid-decode in chunks of this many tokens, with a decode step between chunks. 0 = off.
--max-body-bytes int 0 Cap on request-body size in bytes; a larger body is rejected with 413 before it is read. 0 derives the cap from the model's context window.

Sessions and KV cache

Keeping conversations warm between requests, and across restarts.

Flag Type Default What it does
--kv-sessions int 4 Conversations kept prefilled in RAM for prompt-prefix reuse. 0 disables. On GPU backends this is also how many resident KV slots a model keeps, clamped by its memory guard. Metal keeps 2 slots unless this flag is given (each slot's KV is resident from the first token on unified memory).
--session-dir string — Directory that persists and restores KV sessions across restarts.
--kv-idle-demote duration 0s Demote a warm session's KV to --session-dir once it has been idle this long (for example 10m); it faults back in on the next matching request. 0 = off.
--kv-demoted-max int 64 Maximum demoted (on-disk) sessions to keep; older ones are dropped. Only with --kv-idle-demote.
--job-dir string — Directory for a durable job journal: one line per generation state change. After a restart, a job still running at its last recorded change is marked interrupted. Off by default.

Speculative decoding

Lossless drafting: output is identical to plain decoding, only the speed changes.

Flag Type Default What it does
--spec ngram — Lossless n-gram (prompt-lookup) speculative decoding. Wins on copy-heavy traffic such as code edits, RAG and agent loops. Serves one generation at a time unless --spec-adaptive is also set. More.
--spec-adaptive bool false Experimental. With --spec ngram, keeps concurrent generations: one speculates only while it is alone and joins batched decode when others arrive. Needs a resident whose decode can batch. Graded on CUDA: faster than batching alone, but under load slower than batching on chat (0.86x) and slower than plain --spec ngram on copy (0.80x), so not recommended there (docs/measurements/mc4-candidate-cuda-2026-10-01.md).
--drafter string — A pretrained block-drafter directory (DFlash) paired with --model: it proposes a block of tokens per round and the target verifies them in one pass. Lossless; greedy requests only; needs a resident GPU backend. More.

Network and security

Where the server listens and who may talk to it.

Flag Type Default What it does
--addr string 127.0.0.1:8080 Listen address. Loopback by default; bind 0.0.0.0 only with --api-key and TLS, or behind a TLS-terminating proxy.
--api-key string — Shared secret every request must send as Authorization: Bearer <key> or x-api-key. Falls back to $GOINFER_API_KEY. Required to bind anywhere but loopback, and with --allow-admin.
--tls-cert string — PEM certificate file. With --tls-key, serves HTTPS instead of plaintext HTTP.
--tls-key string — PEM private key file, paired with --tls-cert.
--web bool false Serve a local browser UI at / to chat, and to pull and load GGUF files or whole safetensors checkpoints from Hugging Face. Off by default: its pull route starts multi-gigabyte downloads.
--allow-admin bool false Enable /admin/* on the TCP listener: model load and unload, cancel, halt, resume. A deliberate opt-in that requires --api-key. Ignored when --admin-socket is set. More.

Admin and halt

Controlling a running server without restarting it.

Flag Type Default What it does
--admin-socket string — Serve /admin/* on a Unix socket (mode 0600, no key check: file permissions are the auth) instead of TCP. The same binary controls it with status, ls, cancel, halt and resume.
--halt-file string — Poll this path every 250 ms: present halts the server (inference routes return 503 and in-flight generations are cancelled), absent resumes it. A supervisor halts with touch and resumes with rm.
--halt-exit-code int 0 When nonzero, a halt exits the process with this code once every cancelled generation has stopped, for supervisors whose restart policy must not undo a deliberate halt. 0 = a halt never exits.
--unload-drain-wait duration 5s How long POST /admin/models/unload waits for in-flight requests to drain before returning 202.

Reasoning models

Models that think before they answer (Qwen3, Qwen3.5, Gemma 4): whether they are prompted to, and where the thinking goes in the reply.

Flag Type Default What it does
--thinking template | asis | on | off template Default thinking mode for a model whose chat template has a recognised thinking control (Qwen3, Qwen3.5, Gemma 4): template (the default: what the checkpoint's own template renders), asis (the prompt serve rendered before thinking was modelled: nothing written, the model decides), on, or off. A request overrides it.
--reasoning-format deepseek | deepseek-legacy | none deepseek How a reply's reasoning reaches the client: deepseek (clean content, reasoning in reasoning_content / Anthropic thinking blocks), deepseek-legacy (both, tags kept in content), or none (nothing separated).
--reasoning-budget auto | unlimited | N auto A ceiling on how long a thinking reply may think before serve forces the block closed, so the reply always has room to answer: auto (the default: at most three quarters of the request's max_tokens), unlimited, or N tokens. A request's thinking_token_budget / Anthropic budget_tokens applies too. Off under -spec / -drafter.
--tool-format auto | hermes | template auto How tools are put in the prompt for a model whose own template declares a call form goinfer can render exactly (Qwen3.5's XML, Gemma 4's canonical template): auto (the default: each family's measured default — the model's own form for Gemma 4, goinfer's prompt for Qwen3.5), hermes (goinfer's own prompt) or template (the model's own; a tool_choice naming a function is then a 400).