iPhone / iPad · 2021 · A15

Local AI on the iPhone 13

4 GB is tight. Qwen 3 0.6B or Llama 3.2 1B only. Bigger models belong on a PC via Connect.

Advertised RAM
4 GB
Usable for a model
~1.7 GB
Compute
Metal / ANE
LM Mini path
MLX / Metal

Start in LM Mini

Install the app, pick one of these, chat offline.

Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.

What actually fits

Q4 weights plus KV cache. “Tight” means close Chrome first.

ModelParamsQ4HereWhere
Qwen 3 0.6B 0.6B ~400–500 MB Fits In LM Mini
Qwen 3.5 0.8B 0.8B ~500–600 MB Fits Studio / Ollama
Gemma 3 1B Instruct 1B ~700–750 MB Tight In LM Mini
Llama 3.2 1B Instruct 1B ~700–808 MB Tight In LM Mini

Too big for this phone

Host these on LM Studio or Ollama, then chat from iPhone 13.

How we score this

iPhone 13 ships with 4 GB. After iPhone / iPad we budget ~1.7 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.