Models / Phi-3 / Phi-4generated from capability-matrix.json
Phi-3 / Phi-4
Microsoft's Phi-3 and Phi-4.
chatsafetensors, GGUF
Phi-3 mini 4k
Get it
fits on the CPU on MacBook Pro, M1 Pro, 16 GBfits on the CPU on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull phi3-mini-4k- repo
- microsoft/Phi-3-mini-4k-instruct-gguf
- file
- Phi-3-mini-4k-instruct-q4.gguf
- quant
- q4
- size
- 2.4 GB
- sha256
- 8a83c7fb9049a9b2…
Good for, and what it needs
General instruction following at a size that still loads in seconds.
~3 GB resident at --quant int4
Tool calling
not yet measured
Decisions · /v1/systemone
Label scoring, unmeasured.
What hasn't been shown
- Runs on the CPU only today. Not eligible to live on a GPU.
- Tool calling hasn't been measured on Phi-3 mini 4k.
- No graded speed on the M1 Pro 16 GB.
- No graded speed on the Ryzen 7 3700X.
- Never measured on a machine like yours, so no speed is shown.
How sure we are
Against the released model100.0%
picks the same next token as HuggingFace
1.00000
closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)
full-oracle 100.0%/1.00000 · what parity-gated means
Measured speed · decode, tokens per second
your pick · MacBook Pro, M1 Pro, 16 GBMetal
not measured
your pick · RTX 2070 SUPER, 8 GBCUDA · 2026-09-29
0.90× Ollama At its default, --quant q4k. The earlier 1.14× compared a raw completion with broken output and is withdrawn.
your pick · Ryzen 7 3700X, CPU onlyCPU
not measured
Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.
Architecture, for the curious
- model_type
- phi3
- design
- softmax-GQA
- experts
- dense
- attention window
- none
- QK-norm
- no
- RoPE
- full
- norm
- RMSNorm, pre-norm
- activation
- SwiGLU
- tied head
- no
- modality
- text
- GPU-resident
- no
registry description: Microsoft Phi-3/Phi-4 dense (fused qkv/gate-up, partial rotary)