Models / Qwen3.5-MoEgenerated from capability-matrix.json

Qwen3.5-MoE

Qwen3.5 and 3.6 mixture-of-experts on hybrid attention.

chatexperts: sparse +sharedsafetensors, GGUF

Get it

No vetted checkpoint yet. You can pull any GGUF of this family by repo and quant. goinfer checks it fits before downloading, but nobody here has timed it.

goinfer-chat pull <owner>/<repo>:<quant>

What hasn't been shown

  • Picks the same next token as the reference at 77.5% of positions, not all of them.
  • No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.

How sure we are

Against the released model
77.5% picks the same next token as HuggingFace
0.99069 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

full-oracle 77.5%/0.99069 · what parity-gated means

Measured speed · decode, tokens per second

not measured — no vetted checkpoint to measure

Architecture, for the curious
model_type
qwen3_5_moe, qwen3_5_moe_text
design
gated-linear hybrid (Gated DeltaNet)
experts
sparse +shared
attention window
none
QK-norm
yes
RoPE
partial
norm
RMSNorm, pre-norm
activation
SwiGLU
tied head
no
modality
text
GPU-resident
eligible
registry description: Qwen3.5/3.6 hybrid: Gated DeltaNet + softmax + MoE