Download

One file. Run it.

goinfer v0.22.0, released 2026-10-06. Six platforms, each with its checksum below. Each is a single file, and there is nothing to install.

Release notes on GitHub

Pick a file

Your platform is marked when the browser reports it. If you want a model included, take the 0.5B; if you have a GGUF, take the runtime.

Server goinfer-serve

The server: OpenAI and Anthropic APIs, the web UI, GPU support built in.

PlatformFileSizesha256
macOS · Apple silicon goinfer-serve-darwin-arm6417 MB442e87d4eb0d5e92029a2b09d6a3552cdad9755c95b4e44fa63b3d1242888caa
macOS · Intel goinfer-serve-darwin-amd6418 MB89c6d22f4f11dd811d71c43cf3c294f067178a350d6bd90bacc218e503070407
Linux · x86-64 goinfer-serve-linux-amd6423 MB5a7c4c4c9fc987f288ced3d4af2825c44839c1d4924aa06d394e2be9bd642e46
Linux · ARM64 goinfer-serve-linux-arm6421 MB6d7dc02cbc574ff9e9280e6fd51d486aab875a7933438d408023c86c59c6d76b
Windows · x86-64 goinfer-serve-windows-amd64.exe17 MBe580c9e1cf4c7986c0c32218b3c1370d52e3c6a7d758ea098541376c6eaca44c
Windows · ARM64 goinfer-serve-windows-arm64.exe16 MBcd1499265054d064bccf9128e87f50d818be9ab9255b0cdaf23182b38f5805f8

Chat runtime goinfer-chat

The single-shot runtime. Point it at your own GGUF.

PlatformFileSizesha256
macOS · Apple silicon goinfer-chat-darwin-arm6414 MBefd81073f0a0f1963395b06362d105c59b4c7a2219f979fa975b171377fd69df
macOS · Intel goinfer-chat-darwin-amd6415 MB9bd5ab5296893ff76a182f97bd427848bed246b90d2f15ef8297846da5819fce
Linux · x86-64 goinfer-chat-linux-amd6420 MB7be35048f131d168719e14c7e28d4edd139de9669391cd480f7646c939fb0614
Linux · ARM64 goinfer-chat-linux-arm6419 MB6e7a37e848cca7cc8872b43ae46027fe952bbc9991c0dc5e2d2ddb9c979183eb
Windows · x86-64 goinfer-chat-windows-amd64.exe14 MB74889c24f12e9a987d003a836f1f13d8fd0143bc514e88eaa47467af5559d6b4
Windows · ARM64 goinfer-chat-windows-arm64.exe13 MB70cb6bf121bac90e822b24de8b87c3b90a7dafc9a617e77e3af838859aaa011b

Chat with a 0.5B model inside goinfer-chat-0.5b

Runtime and model in one file. Nothing to install or download.

PlatformFileSizesha256
macOS · Apple silicon goinfer-chat-0.5b-darwin-arm64628 MB2e4220793b1dfe6f11121c0400f9f26afd9fa0cf247f8f74a064eda146c926a1
macOS · Intel goinfer-chat-0.5b-darwin-amd64624 MBa575683dd621c300d8b1217e5607e447c6fe22fedba19883a6eb439ae6c8294a
Linux · x86-64 goinfer-chat-0.5b-linux-amd64629 MBf174a004626b63f53b499735fe14861be64378dc5cd78914b57b3be757976a9b
Linux · ARM64 goinfer-chat-0.5b-linux-arm64628 MB1a302364846ea0ac49ec416c257ca441ce69fbc40179f0719b52dbdb33f7944e
Windows · x86-64 goinfer-chat-0.5b-windows-amd64.exe623 MB12a4c234699149d7883e0ab8a88308b7bc12d4071ca48552e0caebfd6ae3f347
Windows · ARM64 goinfer-chat-0.5b-windows-arm64.exe622 MB141734fc4705fb7523f42c85e28ddc58557ddcae4e8599d5b8f65ae6ee926946

Chat with a 1.5B model inside goinfer-chat-1.5b

The same, with the 1.5B coder model.

PlatformFileSizesha256
macOS · Apple silicon goinfer-chat-1.5b-darwin-arm641.69 GBc4167f72b1e782e89651bd5bf03d03465bdbb1c24df31bdace06a96e1f2f518e
macOS · Intel goinfer-chat-1.5b-darwin-amd641.68 GBf614051d3aa406b13b421d51746f995a1bfb850b1d5c6e148101fbda15bf0bd2
Linux · x86-64 goinfer-chat-1.5b-linux-amd641.68 GB11ec5ffdf3069686d9dbc3a3c3ad200fc22d55d27acc5028f7f2e1e3e5fb5e38
Linux · ARM64 goinfer-chat-1.5b-linux-arm641.68 GBe7912db52897a9a85909da5351fa90dd609b1672dcf3a10896fce948b2ddb145
Windows · x86-64 goinfer-chat-1.5b-windows-amd64.exe1.68 GBc0f606f99bf9e5ac5e9e81d5fa67cf9f01f77427a09b86e29a61a1bec12bbc03
Windows · ARM64 goinfer-chat-1.5b-windows-arm64.exe1.68 GB8031cee8f925fc0f1d97e89b2fba62af95207cc933814e0282e78874daf262a0

Run it, and check it

chmod +x goinfer-chat-0.5b-darwin-arm64
./goinfer-chat-0.5b-darwin-arm64

The same commands work for the server and the runtime. To bring your own model, give the runtime a GGUF with --model, or let it fetch one: ./goinfer-chat-<os>-<arch> pull qwen2.5-coder-0.5b. To check a download against the release's own list:

grep goinfer-chat-0.5b-darwin-arm64 checksums.txt | shasum -a 256 -c

Every binary's sha256 is in checksums.txt. Other files on the release: NOTICE.txt QWEN2.5-CODER-LICENSE.txt .

What it has been run on

goinfer is built, measured and checked on MacBook Pro · M1 Pro 16 GB, RTX 2070 SUPER · 8 GB, and Ryzen 7 · CPU only. Every speed on this site came from one of those, and does not carry over to other hardware. Other machines, other GPU generations and other operating systems have mostly not been run by us, and what is and is not listed as having run is in the generated hardware matrix.

What stands between you and a wrong answer on a machine we have not seen is a check goinfer runs on your machine at start: it runs the compute kernels it is about to use against a reference on a small fixed input, and if one disagrees it steps down to a slower path, or declines that backend, rather than giving wrong numbers. Today that covers the CPU kernels, CUDA, WebGPU on a real GPU (not on a software renderer), and Metal; each GPU backend's margins were measured on one device so far, so a healthy GPU of another kind could be declined to the slower path. -no-selftest turns it off, and goinfer-serve check --hardware (or goinfer-chat check --hardware) prints what it found on yours, with nothing sent anywhere.

If something is wrong on your hardware, a bug report with that output is the most useful thing you can send.

As a Go library

go get github.com/townsendmerino/goinfer/decoder@latest github.com/townsendmerino/goinfer/tokenizer@latest

Run it from inside your own module, and name the packages you import: a bare go get of the module records the requirement without fetching enough to build, and the next build fails with missing go.sum entry. One such command covers every other package in the module. The engine has no cgo on the default build and needs Go 1.27 or newer. Its README has the API and a complete forty-line program.