Models / gpt-ossgenerated from capability-matrix.json
gpt-oss
OpenAI's open-weight gpt-oss, 20B and 120B.
chatexperts: sparse, no-sharedsafetensors, GGUF
gpt-oss 20B
Get it
tight pages experts from disk on MacBook Pro, M1 Pro, 16 GBfits experts streamed to the card on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull gpt-oss-20b- repo
- ggml-org/gpt-oss-20b-GGUF
- file
- gpt-oss-20b-MXFP4.gguf
- quant
- mxfp4
- size
- 12.1 GB
- sha256
- 27cd6c432c7672cb…
Good for, and what it needs
The 20-35B-class MoE this project actually validates and measures: too big to hold fully resident on an 8 GB GPU, which is the point — bring -moe-cache-experts or -stream-weights.
~12 GB if loaded fully resident (native MXFP4); on an 8 GB card use -moe-cache-experts (CUDA resident-core + cached-experts, measured working, not just eligible — see this family's description above)
Tool calling
not yet measured
Decisions · /v1/systemone
Label scoring, unmeasured.
What hasn't been shown
- Tool calling hasn't been measured on gpt-oss 20B.
- No speed measured on any machine.
- No graded speed on the M1 Pro 16 GB.
- No graded speed on the RTX 2070 SUPER.
- No graded speed on the Ryzen 7 3700X.
- Never measured on a machine like yours, so no speed is shown.
How sure we are
Against the released model100.0%
picks the same next token as HuggingFace
0.99843
closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)
real-oracle 100.0%/0.99843 · what parity-gated means
Measured speed · decode, tokens per second
your pick · MacBook Pro, M1 Pro, 16 GBMetal
not measured
your pick · RTX 2070 SUPER, 8 GBCUDA
not measured
your pick · Ryzen 7 3700X, CPU onlyCPU
not measured
Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.
Architecture, for the curious
- model_type
- gpt_oss
- design
- softmax-GQA
- experts
- sparse, no-shared
- attention window
- interleave
- QK-norm
- no
- RoPE
- full
- norm
- RMSNorm, pre-norm
- activation
- SwiGLU
- tied head
- no
- modality
- text
- GPU-resident
- eligible
registry description: OpenAI gpt-oss 20b/120b sparse MoE: per-head attention sinks + clamped interleaved-SwiGLU + alternating sliding/full + YaRN (MXFP4 experts; GPU-resident on BOTH Metal and CUDA since 2026-08-31 — the CUDA half validated on the real 20B, resident on an 8 GB card via --moe-cache-experts)