iPhone / iPad · 2021 · A15
Local AI on the iPhone 13
4 GB is tight. Qwen 3 0.6B or Llama 3.2 1B only. Bigger models belong on a PC via Connect.
- Advertised RAM
- 4 GB
- Usable for a model
- ~1.7 GB
- Compute
- Metal / ANE
- LM Mini path
- MLX / Metal
Start in LM Mini
Install the app, pick one of these, chat offline.
Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.
What actually fits
Q4 weights plus KV cache. “Tight” means close Chrome first.
| Model | Params | Q4 | Here | Where |
|---|---|---|---|---|
| Qwen 3 0.6B | 0.6B | ~400–500 MB | Fits | In LM Mini |
| Qwen 3.5 0.8B | 0.8B | ~500–600 MB | Fits | Studio / Ollama |
| Gemma 3 1B Instruct | 1B | ~700–750 MB | Tight | In LM Mini |
| Llama 3.2 1B Instruct | 1B | ~700–808 MB | Tight | In LM Mini |
Too big for this phone
Host these on LM Studio or Ollama, then chat from iPhone 13.
- DeepSeek R1 Distill 1.5B — PC + Connect
- Qwen 3 1.7B — PC + Connect
- Llama 3.2 3B Instruct — PC + Connect
- Phi-4 Mini Instruct — PC + Connect
- Qwen 3 4B Instruct 2507 — PC + Connect
- Gemma 3 4B Instruct — PC + Connect
How we score this
iPhone 13 ships with 4 GB. After iPhone / iPad we budget ~1.7 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.