iPhone / iPad · 2024 · A18 Pro
Local AI on the iPhone 16 Pro
Same 8 GB and A18 Pro as the Max. Local AI feels snappy on 1.7B–4B; leave 8B for a Mac or Connect.
- Advertised RAM
- 8 GB
- Usable for a model
- ~3.6 GB
- Compute
- Metal / ANE
- LM Mini path
- MLX / Metal
Start in LM Mini
Install the app, pick one of these, chat offline.
Strong reasoning in a phone-friendly package — great for explanations and code-ish help.
Llama 3.2 3B Instruct Fits · ~1.9–2.0 GBNoticeably smarter than 1B — good for writing help and longer chats on Pro phones.
Qwen 3 1.7B Runs well · ~1.1–1.2 GBBest everyday phone pick for most people — chat, tools, and general help.
DeepSeek R1 Distill 1.5B Runs well · ~1.0–1.1 GBA mini “thinking” model — stronger at math and step-by-step reasoning.
Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.
What actually fits
Q4 weights plus KV cache. “Tight” means close Chrome first.
| Model | Params | Q4 | Here | Where |
|---|---|---|---|---|
| Qwen 3 0.6B | 0.6B | ~400–500 MB | Runs well | In LM Mini |
| Gemma 3 1B Instruct | 1B | ~700–750 MB | Runs well | In LM Mini |
| Llama 3.2 1B Instruct | 1B | ~700–808 MB | Runs well | In LM Mini |
| DeepSeek R1 Distill 1.5B | 1.5B | ~1.0–1.1 GB | Runs well | In LM Mini |
| Qwen 3 1.7B | 1.7B | ~1.1–1.2 GB | Runs well | In LM Mini |
| Qwen 3.5 0.8B | 0.8B | ~500–600 MB | Runs well | Studio / Ollama |
| SmolLM2 1.7B Instruct | 1.7B | ~1.0 GB | Runs well | Studio / Ollama |
| Qwen 2.5 1.5B Instruct | 1.5B | ~1.0 GB | Runs well | Studio / Ollama |
| Llama 3.2 3B Instruct | 3B | ~1.9–2.0 GB | Fits | In LM Mini |
| Phi-4 Mini Instruct | 3.8B | ~2.3–2.4 GB | Fits | In LM Mini |
| Qwen 3.5 2B | 2B | ~1.3–1.5 GB | Fits | Studio / Ollama |
| Gemma 3n E2B | 2B | ~1.5 GB | Fits | Studio / Ollama |
| Gemma 4 E2B Instruct | 2.3B | ~1.4–1.8 GB (Q4/QAT) | Fits | Studio / Ollama |
| Qwen 2.5 3B Instruct | 3B | ~1.9 GB | Fits | Studio / Ollama |
| SmolLM3 3B | 3B | ~1.9 GB (Q4) | Fits | Studio / Ollama |
| Qwen 3 4B Instruct 2507 | 4B | ~2.4–2.5 GB | Tight | In LM Mini |
| Nemotron 3 Nano 4B | 4B | ~2.5 GB (Q4) | Tight | Studio / Ollama |
| Qwen 3.5 4B | 4B | ~2.5–2.8 GB | Tight | Studio / Ollama |
| Gemma 3n E4B | 4B | ~2.6 GB (Q4) | Tight | Studio / Ollama |
Too big for this phone
Host these on LM Studio or Ollama, then chat from iPhone 16 Pro.
- Gemma 3 4B Instruct — PC + Connect
- Qwen 3 8B — PC + Connect
- Gemma 3 12B Instruct — PC + Connect
- Qwen 3 14B — PC + Connect
- Gemma 3 27B Instruct — PC + Connect
- Qwen 3 32B — PC + Connect
How we score this
iPhone 16 Pro ships with 8 GB. After iPhone / iPad we budget ~3.6 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.