Models / Ling 3.0generated from capability-matrix.json

Ling 3.0

inclusionAI's Ling 3.0, tiny and flash.

chatexperts: sparse +sharedexperimentalsafetensors

Get it

goinfer can't fetch this one yet. Download the safetensors folder with your usual tool, then point goinfer at it.

goinfer-serve -model ./path/to/checkpoint-folder

What hasn't been shown

  • Only checked against a small test model built from the same wiring, because no released Ling 3.0 was small enough to run on the hardware available. Two families promoted past this stage turned out to have real bugs behind a passing fixture.
  • Registered as experimental. It can change, or go, before v1.0.
  • Runs on the CPU only today. Not eligible to live on a GPU.
  • goinfer can't download it for you: this family only loads from safetensors, which come as several files. That's planned (checkpoint fetch, P1–P9).
  • No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.

How sure we are

Against a small test model
100.0% picks the same next token as HuggingFace
1.00000 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

experimental: tiny-oracle 100.0%/1.00000 · what parity-gated means

Measured speed · decode, tokens per second

not measured — no vetted checkpoint to measure

Architecture, for the curious
model_type
bailing_hybrid
design
latent-KV (MLA)
experts
sparse +shared
attention window
none
QK-norm
no
RoPE
full
norm
RMSNorm, pre-norm
activation
SwiGLU
tied head
no
modality
text
GPU-resident
no
registry description: inclusionAI Ling 3.0 (tiny/flash): DeepSeek-style MLA alternating with Kimi Delta Attention (per-channel-decay delta rule) every layer_group_size-th layer, over a DeepSeekMoE FFN