// model_catalog
Curated the hard way: we ran 29 candidates through the same on-device tests and kept the 12 that earned a place.
Every one is quantized for Apple Silicon and runs fully on-device. Grouped by where it runs, then by family — memory figures are on-device peaks, weights are download size.
// how_this_was_measured
Measured in Mortise against 72 deterministic tasks spanning factual recall, arithmetic, code reading, logic, instruction-following and extraction. Models run as 4-bit MLX conversions (3-bit for some), under this app's prompts, sampling and memory limits.
This describes how each conversion behaves inside Mortise. It is not a general benchmark of the underlying model, and quantized conversions do not represent the full-precision weights their makers published.
Models that fit within iOS's ~5 GB per-app memory budget. “Tight” models still run, but sit close to the limit on some devices.
Ultra-light chat and quick PDF Q&A; long answers cut off
Fast reasoning, PDF Q&A, and doc lookup at a tiny size
Reasoning plus accurate doc lookup in a light model
Top all-rounder — reasoning and PDF tools; passed every test
Fast, dependable chat and PDF/doc Q&A (Liquid AI)
Reasoning, PDFs, and doc lookup — passed every on-device test
Reasoning and doc Q&A — near-perfect in tests
Compact instruction-following chat and doc Q&A
Every catalog model runs on a Mac with enough memory — the card shows the minimum RAM each one needs.
Ultra-light chat and quick PDF Q&A; long answers cut off
Fast reasoning, PDF Q&A, and doc lookup at a tiny size
Reasoning plus accurate doc lookup in a light model
Top all-rounder — reasoning and PDF tools; passed every test
Strongest reasoning in class — passed every on-device test
Fast, dependable chat and PDF/doc Q&A (Liquid AI)
Reasoning, PDFs, and doc lookup — passed every on-device test
Reasoning and doc Q&A — near-perfect in tests
Long-context reasoning and doc Q&A on Mac
Reasoning and docs — passed every on-device test
Compact instruction-following chat and doc Q&A
Standings describe how each conversion behaves inside Mortise — not a general benchmark of the underlying model. How this was measured
* Context is the developer’s advertised maximum, not the window Mortise will give you. The app calculates that at runtime — up to the listed figure, if your hardware has the headroom.
// buy_once
No subscription, no account, no cloud. Every update included.