Models / Phi-3 / Phi-4generated from capability-matrix.json

Phi-3 / Phi-4

Microsoft's Phi-3 and Phi-4.

chatsafetensors, GGUF

Phi-3 mini 4k

Get it

fits on the CPU on MacBook Pro, M1 Pro, 16 GBfits on the CPU on RTX 2070 SUPER, 8 GBfits on Ryzen 7 3700X, CPU only
goinfer-chat pull phi3-mini-4k
repo
microsoft/Phi-3-mini-4k-instruct-gguf
file
Phi-3-mini-4k-instruct-q4.gguf
quant
q4
size
2.4 GB
sha256
8a83c7fb9049a9b2…

Good for, and what it needs

General instruction following at a size that still loads in seconds.

~3 GB resident at --quant int4

Tool calling

not yet measured

Decisions · /v1/systemone

Label scoring, unmeasured.

What hasn't been shown

  • Runs on the CPU only today. Not eligible to live on a GPU.
  • Tool calling hasn't been measured on Phi-3 mini 4k.
  • No graded speed on the M1 Pro 16 GB.
  • No graded speed on the Ryzen 7 3700X.
  • Never measured on a machine like yours, so no speed is shown.

How sure we are

Against the released model
100.0% picks the same next token as HuggingFace
1.00000 closest the raw scores get at their worst position (cosine, 1.00000 is identical; bar starts at 0.95)

full-oracle 100.0%/1.00000 · what parity-gated means

Measured speed · decode, tokens per second

your pick · MacBook Pro, M1 Pro, 16 GBMetal
not measured
your pick · RTX 2070 SUPER, 8 GBCUDA · 2026-09-29
goinfer
112.6 Ollama
125.8
0.90× Ollama At its default, --quant q4k. The earlier 1.14× compared a raw completion with broken output and is withdrawn.
your pick · Ryzen 7 3700X, CPU onlyCPU
not measured

Ollama v0.32.5 at its defaults, same machine, same session. 128 tokens of context unless noted. Full method in benchmarks.

Architecture, for the curious
model_type
phi3
design
softmax-GQA
experts
dense
attention window
none
QK-norm
no
RoPE
full
norm
RMSNorm, pre-norm
activation
SwiGLU
tied head
no
modality
text
GPU-resident
no
registry description: Microsoft Phi-3/Phi-4 dense (fused qkv/gate-up, partial rotary)