Models / Qwen3.5-MoEgenerated from capability-matrix.json
Qwen3.5-MoE
Qwen3.5 and 3.6 mixture-of-experts on hybrid attention.
chatexperts: sparse +sharedsafetensors, GGUF
Get it
No vetted checkpoint yet. You can pull any GGUF of this family by repo and quant. goinfer checks it fits before downloading, but nobody here has timed it.
goinfer-chat pull <owner>/<repo>:<quant>What hasn't been shown
- Picks the same next token as the reference at 77.5% of positions, not all of them.
- No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.
How sure we are
Against the released model77.5%
picks the same next token as HuggingFace
0.99069
closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)
full-oracle 77.5%/0.99069 · what parity-gated means
Measured speed · decode, tokens per second
not measured — no vetted checkpoint to measure
Architecture, for the curious
- model_type
- qwen3_5_moe, qwen3_5_moe_text
- design
- gated-linear hybrid (Gated DeltaNet)
- experts
- sparse +shared
- attention window
- none
- QK-norm
- yes
- RoPE
- partial
- norm
- RMSNorm, pre-norm
- activation
- SwiGLU
- tied head
- no
- modality
- text
- GPU-resident
- eligible
registry description: Qwen3.5/3.6 hybrid: Gated DeltaNet + softmax + MoE