Android · 2025 · Snapdragon 8 Elite
Local AI on the Galaxy S25 Ultra
2025 Ultra. Adreno + 12 GB is a strong Vulkan target — 4B easy, 8B possible, 14B is a PC job.
- Advertised RAM
- 12 GB
- Usable for a model
- ~6.2 GB
- Compute
- Vulkan
- LM Mini path
- GGUF / Vulkan
Start in LM Mini
Install the app, pick one of these, chat offline.
Balanced daily driver for Pro phones and Macs — quality without a desktop GPU.
Gemma 3 4B Instruct Runs well · ~2.7–3.7 GBGoogle’s mid-size Gemma 3 — sharper answers; some builds can look at images.
Phi-4 Mini Instruct Runs well · ~2.3–2.4 GBStrong reasoning in a phone-friendly package — great for explanations and code-ish help.
Llama 3.2 3B Instruct Runs well · ~1.9–2.0 GBNoticeably smarter than 1B — good for writing help and longer chats on Pro phones.
Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.
What actually fits
Q4 weights plus KV cache. “Tight” means close Chrome first.
| Model | Params | Q4 | Here | Where |
|---|---|---|---|---|
| Qwen 3 0.6B | 0.6B | ~400–500 MB | Runs well | In LM Mini |
| Gemma 3 1B Instruct | 1B | ~700–750 MB | Runs well | In LM Mini |
| Llama 3.2 1B Instruct | 1B | ~700–808 MB | Runs well | In LM Mini |
| DeepSeek R1 Distill 1.5B | 1.5B | ~1.0–1.1 GB | Runs well | In LM Mini |
| Qwen 3 1.7B | 1.7B | ~1.1–1.2 GB | Runs well | In LM Mini |
| Llama 3.2 3B Instruct | 3B | ~1.9–2.0 GB | Runs well | In LM Mini |
| Phi-4 Mini Instruct | 3.8B | ~2.3–2.4 GB | Runs well | In LM Mini |
| Qwen 3 4B Instruct 2507 | 4B | ~2.4–2.5 GB | Runs well | In LM Mini |
| Gemma 3 4B Instruct | 4B | ~2.7–3.7 GB | Runs well | In LM Mini |
| Qwen 3.5 0.8B | 0.8B | ~500–600 MB | Runs well | Studio / Ollama |
| SmolLM2 1.7B Instruct | 1.7B | ~1.0 GB | Runs well | Studio / Ollama |
| Qwen 2.5 1.5B Instruct | 1.5B | ~1.0 GB | Runs well | Studio / Ollama |
| Qwen 3.5 2B | 2B | ~1.3–1.5 GB | Runs well | Studio / Ollama |
| Gemma 3n E2B | 2B | ~1.5 GB | Runs well | Studio / Ollama |
| Gemma 4 E2B Instruct | 2.3B | ~1.4–1.8 GB (Q4/QAT) | Runs well | Studio / Ollama |
| Qwen 2.5 3B Instruct | 3B | ~1.9 GB | Runs well | Studio / Ollama |
| SmolLM3 3B | 3B | ~1.9 GB (Q4) | Runs well | Studio / Ollama |
| Nemotron 3 Nano 4B | 4B | ~2.5 GB (Q4) | Runs well | Studio / Ollama |
| Qwen 3.5 4B | 4B | ~2.5–2.8 GB | Runs well | Studio / Ollama |
| Gemma 3n E4B | 4B | ~2.6 GB (Q4) | Runs well | Studio / Ollama |
| Gemma 4 E4B Instruct | 4.5B | ~2.6–3.2 GB (Q4/QAT) | Runs well | Studio / Ollama |
| Mistral 7B Instruct | 7B | ~4.1 GB (Q4) | Tight | Studio / Ollama |
| OLMo 3 7B Instruct | 7B | ~4.3 GB (Q4) | Tight | Studio / Ollama |
| Qwen 2.5 7B Instruct | 7B | ~4.4 GB (Q4) | Tight | Studio / Ollama |
Too big for this phone
Host these on LM Studio or Ollama, then chat from Galaxy S25 Ultra.
- Qwen 3 8B — PC + Connect
- Gemma 3 12B Instruct — PC + Connect
- Qwen 3 14B — PC + Connect
- Gemma 3 27B Instruct — PC + Connect
- Qwen 3 32B — PC + Connect
- Llama 3.1 8B Instruct — PC + Connect
How we score this
Galaxy S25 Ultra ships with 12 GB. After Android we budget ~6.2 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.