Models / GLM-4.5/4.6generated from capability-matrix.json
GLM-4.5/4.6
Zhipu's GLM-4.5 and 4.6 mixture-of-experts.
chatexperts: sparse +sharedexperimentalsafetensors, GGUF
Get it
No vetted checkpoint yet. You can pull any GGUF of this family by repo and quant. goinfer checks it fits before downloading, but nobody here has timed it.
goinfer-chat pull <owner>/<repo>:<quant>What hasn't been shown
- Only checked against a small test model built from the same wiring, because no released GLM-4.5/4.6 was small enough to run on the hardware available. Two families promoted past this stage turned out to have real bugs behind a passing fixture.
- Registered as experimental. It can change, or go, before v1.0.
- No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.
How sure we are
Against a small test model100.0%
picks the same next token as HuggingFace
1.00000
closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)
experimental: tiny-oracle 100.0%/1.00000 · what parity-gated means
Measured speed · decode, tokens per second
not measured — no vetted checkpoint to measure
Architecture, for the curious
- model_type
- glm4_moe
- design
- softmax-GQA
- experts
- sparse +shared
- attention window
- none
- QK-norm
- yes
- RoPE
- partial
- norm
- RMSNorm, pre-norm
- activation
- SwiGLU
- tied head
- no
- modality
- text
- GPU-resident
- eligible
registry description: Zhipu GLM-4.5/4.6 DeepSeek-style MoE (sigmoid routing + dense prefix)