Models / GLM-OCRgenerated from capability-matrix.json
GLM-OCR
Zhipu's GLM-OCR document model. Reads an image of a page: text, a table or a formula out, or, given a schema, JSON that matches it (goinfer-chat --image invoice.png --schema invoice.schema.json). Verified on the real checkpoint, token-identical to transformers at f32. Images up to 6,144 image tokens, about 4.8 megapixels (the default); --vision-max-pixels lowers the cap, and the CPU vision tower's cost falls with it.
chatimagesexperimentalsafetensors
Get it
goinfer can't fetch this one yet. Download the safetensors folder with your usual tool, then point goinfer at it.
goinfer-serve -model ./path/to/checkpoint-folderWhat hasn't been shown
- Only checked against a small test model built from the same wiring, because no released GLM-OCR was small enough to run on the hardware available. Two families promoted past this stage turned out to have real bugs behind a passing fixture.
- Registered as experimental. It can change, or go, before v1.0.
- goinfer can't download it for you: this family only loads from safetensors, which come as several files. That's planned (checkpoint fetch, P1–P9).
- No vetted checkpoint, so there's no fit verdict, no speed and no tool-calling result for this family.
How sure we are
Against a small test model100.0%
picks the same next token as HuggingFace
1.00000
closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)
experimental: tiny-oracle 100.0%/1.00000 · what parity-gated means
Measured speed · decode, tokens per second
not measured — no vetted checkpoint to measure
Architecture, for the curious
- model_type
- glm_ocr, glm_ocr_text
- design
- softmax-GQA
- experts
- dense
- attention window
- none
- QK-norm
- no
- RoPE
- m-RoPE
- norm
- RMSNorm, sandwich
- activation
- SwiGLU
- tied head
- no
- modality
- text (+ vision tower, CPU)
- GPU-resident
- eligible
registry description: Zhipu GLM-OCR (0.9B document OCR): GLM-4V's 4-norm sandwich with GLM's own tensor names, fused gate_up, explicit head_dim, pairwise m-RoPE with contiguous sections; MTP layer skipped; its own vision tower (aikit, CPU) feeds merged rows through the Qwen-shaped image path