Model guide

Which model should I run?

A plain-language list of 51 local models. Read the first sentence if you are new; the grey line is for specs. Download smaller ones in LM Mini, or load bigger ones on a home PC with LM Studio / Ollama and connect remotely. Not sure about RAM? Find your phone.

Phone free RAM needed on the handset Mac M-series unified memory GPU desktop VRAM if you host at home

Everyday phones

Small models that fit most iPhones and Androids. Start here if you are new to local AI.

  1. Qwen 3 0.6B

    In LM Mini

    Tiny and quick — good for short replies when your phone is low on memory.

    0.6B ~400–500 MB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    Qwen3-0.6B; Q4 GGUF & 4-bit MLX. Hugging Face downloads in the tens of millions.

  2. Qwen 3.5 0.8B

    LM Studio / Ollama

    2026 Qwen tiny — hybrid long-context brain in a phone-sized file. Best when you want the newest Qwen on 4–6 GB devices.

    0.8B ~500–600 MB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    Qwen3.5-0.8B (Feb 2026). Gated DeltaNet hybrid; 262K native context. Community GGUF/MLX as engines catch up — import in LM Mini or host via Connect.

  3. Gemma 3 1B Instruct

    In LM Mini

    Google’s smallest Gemma 3 chat model. Great starter for everyday questions.

    1B ~700–750 MB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    google/gemma-3-1b-it. Instruct-tuned; GGUF + MLX.

  4. Gemma 4 E2B Instruct

    LM Studio / Ollama

    Google’s 2026 edge Gemma — ~2.3B effective, text + image + audio. The new tiny phone default if your engine supports Gemma 4.

    2.3B ~1.4–1.8 GB (Q4/QAT) Phone 4 GB+ Mac 8 GB+ GPU Any GGUF

    google/gemma-4-E2B-it (Apr 2026). Official QAT GGUF. Apache 2.0. ~5.1B with embeddings; plan ~2–3 GB live.

  5. Meta’s pocket Llama — fast, tool-friendly, and solid for light chat.

    1B ~700–808 MB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    Llama-3.2-1B-Instruct; tool calling. GGUF + MLX.

  6. SmolLM2 1.7B Instruct

    LM Studio / Ollama

    Hugging Face’s small all-rounder. Surprisingly capable for its size.

    1.7B ~1.0 GB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF

    SmolLM2-1.7B-Instruct Q4_K_M via LM Studio / custom HF import.

  7. A mini “thinking” model — stronger at math and step-by-step reasoning.

    1.5B ~1.0–1.1 GB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    R1 distill of Qwen 2.5 1.5B; reasoning traces. GGUF + MLX.

  8. Qwen 3 1.7B

    In LM Mini

    Best everyday phone pick for most people — chat, tools, and general help.

    1.7B ~1.1–1.2 GB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    Qwen3-1.7B; primary free-slot family in LM Mini. GGUF + MLX.

  9. Qwen 3.5 2B

    LM Studio / Ollama

    Newest small Qwen that still fits everyday phones — sharper than 1.7B-class, still downloadable on cellular.

    2B ~1.3–1.5 GB Phone 4 GB+ Mac 8 GB+ GPU Any GGUF · MLX

    Qwen3.5-2B. Hybrid attention + native multimodal in the 3.5 family. Prefer Q4 GGUF once your llama.cpp/MLX build lists Qwen3.5.

  10. Qwen 2.5 1.5B Instruct

    LM Studio / Ollama

    Previous-gen Qwen that’s still a reliable tiny assistant.

    1.5B ~1.0 GB Phone 3 GB+ Mac 8 GB+ GPU Any GGUF

    Qwen2.5-1.5B-Instruct; still huge Hugging Face download volume.

  11. Gemma 3n E2B

    LM Studio / Ollama

    Google’s on-device Gemma 3n (effective 2B). Built for phones, including audio-aware builds.

    2B ~1.5 GB Phone 4 GB+ Mac 8 GB+ GPU Any GGUF

    google/gemma-3n-E2B-it. LiteRT / GGUF community quants. Newer than Gemma 3 1B.

Pro phones & tablets

More capable answers — needs roughly 6 GB+ free RAM on the device.

  1. Noticeably smarter than 1B — good for writing help and longer chats on Pro phones.

    3B ~1.9–2.0 GB Phone 6 GB+ Mac 8–16 GB GPU 6 GB+ VRAM GGUF · MLX

    Llama-3.2-3B-Instruct; tool calling. GGUF + MLX.

  2. SmolLM3 3B

    LM Studio / Ollama

    Hugging Face’s 2025 small model — a step up from SmolLM2, still phone-sized.

    3B ~1.9 GB (Q4) Phone 6 GB+ Mac 8–16 GB GPU 6 GB+ VRAM GGUF

    HuggingFaceTB/SmolLM3-3B. GGUF via bartowski / community.

  3. Gemma 3 4B Instruct

    In LM Mini

    Google’s mid-size Gemma 3 — sharper answers; some builds can look at images.

    4B ~2.7–3.7 GB Phone 6 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF · MLX

    google/gemma-3-4b-it; multimodal on supported builds. GGUF + MLX.

  4. Qwen 3.5 4B

    LM Studio / Ollama

    2026’s 4B daily driver — long context and stronger tools than Qwen 3 4B, if your phone has ~6 GB free.

    4B ~2.5–2.8 GB Phone 6 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF · MLX

    Qwen3.5-4B; millions of HF downloads. Hybrid Gated DeltaNet. Import a Q4 when the engine supports it; Qwen 3 4B Instruct 2507 remains the in-app pick today.

  5. Gemma 4 E4B Instruct

    LM Studio / Ollama

    Google’s on-device Gemma 4 (~4.5B effective) with image and audio. The 2026 upgrade from Gemma 3 4B / 3n E4B.

    4.5B ~2.6–3.2 GB (Q4/QAT) Phone 8 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF

    google/gemma-4-E4B-it. Official QAT GGUF. ~8B with embeddings; budget ~4 GB live. Apache 2.0.

  6. Phi-4 Mini Instruct

    In LM Mini

    Strong reasoning in a phone-friendly package — great for explanations and code-ish help.

    3.8B ~2.3–2.4 GB Phone 6 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF · MLX

    microsoft/Phi-4-mini-instruct (~3.8B). GGUF + MLX.

  7. Balanced daily driver for Pro phones and Macs — quality without a desktop GPU.

    4B ~2.4–2.5 GB Phone 6 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF · MLX

    July 2025 Instruct refresh; millions of HF downloads. GGUF + MLX.

  8. Qwen 2.5 3B Instruct

    LM Studio / Ollama

    Compact Qwen for chat and light coding when 4B feels heavy.

    3B ~1.9 GB Phone 6 GB+ Mac 8–16 GB GPU 6 GB+ VRAM GGUF

    Qwen2.5-3B-Instruct Q4_K_M; LM Studio / Ollama.

  9. Gemma 3n E4B

    LM Studio / Ollama

    Gemma 3n effective-4B — Google’s newer on-device stack, including multimodal 3n builds.

    4B ~2.6 GB (Q4) Phone 8 GB+ Mac 8–16 GB GPU 8 GB+ VRAM GGUF

    google/gemma-3n-E4B-it. Community GGUF (bartowski). Heavier than E2B.

  10. Nemotron 3 Nano 4B

    LM Studio / Ollama

    NVIDIA’s 2026 tiny Nemotron — dense 4B-class, aimed at local agents.

    4B ~2.5 GB (Q4) Phone 8 GB+ Mac 16 GB+ GPU 8 GB+ VRAM GGUF

    nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Use a GGUF quant on phone; BF16 is a desktop download.

Mac (M-series) & strong phones

8B-class models. Comfortable on 16 GB+ Macs; some high-RAM phones can run them too.

  1. Mistral 7B Instruct

    LM Studio / Ollama

    Classic open model — clear writing and solid general knowledge on Mac or GPU.

    7B ~4.1 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 8 GB+ VRAM GGUF

    Mistral-7B-Instruct v0.3; ubiquitous GGUF quants.

  2. Llama 3.1 8B Instruct

    LM Studio / Ollama

    Meta’s workhorse 8B — excellent general assistant on Mac or a home GPU.

    8B ~4.7 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 8–12 GB VRAM GGUF · MLX

    Llama-3.1-8B-Instruct; tool calling; LM Studio / Ollama staple.

  3. Qwen 3 8B

    In LM Mini

    Frontier feel on-device for high-RAM phones and M-series Macs.

    8B ~4.7–4.9 GB Phone 12 GB+ Mac 16 GB+ GPU 10–12 GB VRAM GGUF · MLX

    Qwen3-8B; tool calling. GGUF + MLX in LM Mini.

  4. Qwen 3.5 9B

    LM Studio / Ollama

    Frontier feel on a 16 GB phone or any M-series Mac — the 2026 8B-class upgrade.

    9B ~5.2–5.6 GB (Q4) Phone 16 GB+ Mac 18–24 GB GPU 12 GB+ VRAM GGUF · MLX

    Qwen3.5-9B (most-downloaded 3.5 size on HF). Not in the LM Mini in-app list yet; LM Studio / Ollama / custom import + Connect.

  5. Qwen 2.5 7B Instruct

    LM Studio / Ollama

    Popular all-rounder for coding help, summaries, and chat on desktop or Mac.

    7B ~4.4 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 8–10 GB VRAM GGUF · MLX

    Qwen2.5-7B-Instruct — still one of the most downloaded instruct models on HF.

  6. DeepSeek R1 Distill 8B

    LM Studio / Ollama

    Reasoning-focused — walks through hard problems before answering.

    8B ~4.7 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 10–12 GB VRAM GGUF

    R1 distill on Llama 8B; longer CoT traces.

  7. OLMo 3 7B Instruct

    LM Studio / Ollama

    Allen AI’s fully open 7B instruct model — a transparent alternative to Llama/Qwen.

    7B ~4.3 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 8–10 GB VRAM GGUF

    allenai/Olmo-3-7B-Instruct (2025). Community GGUF.

  8. Nemotron Nano 9B v2

    LM Studio / Ollama

    NVIDIA 9B local model — stronger than 8B class if you have the RAM.

    9B ~5.2 GB (Q4) Phone 16 GB+ Mac 18–24 GB GPU 12 GB+ VRAM GGUF

    nvidia/NVIDIA-Nemotron-Nano-9B-v2. Prefer Q4 GGUF on Mac/GPU.

  9. Gemma 2 9B Instruct

    LM Studio / Ollama

    Google’s 9B chat model — polished answers when you have Mac RAM or a GPU.

    9B ~5.4 GB (Q4) Phone 16 GB+ Mac 18–24 GB GPU 12 GB+ VRAM GGUF · MLX

    Gemma-2-9B-IT; still excellent instruction following.

  10. Hermes 3 8B

    LM Studio / Ollama

    Community favorite for creative writing and agent-style tool use.

    8B ~4.7 GB (Q4) Phone 12 GB+ Mac 16 GB+ GPU 10–12 GB VRAM GGUF

    NousResearch Hermes-3-Llama-3.1-8B; function calling.

High-memory Mac

12B–14B class. Best on 24–32 GB+ unified memory, or remote from a GPU PC.

  1. Qwen 3 30B-A3B Instruct

    LM Studio / Ollama

    Mixture-of-experts: 30B total, ~3B active. Big-model taste on a strong Mac or GPU.

    30B ~16 GB (Q4 MoE) Phone — Mac 32 GB+ GPU 16–24 GB VRAM GGUF

    Qwen3-30B-A3B-Instruct-2507 MoE. Q4 is lighter than dense 30B; still not a phone model.

  2. Qwen 3.5 27B

    LM Studio / Ollama

    Dense 2026 Qwen for a high-RAM Mac or a 24 GB GPU — then use the phone as the remote keyboard.

    27B ~16 GB (Q4) Phone — Mac 32 GB+ GPU 24 GB+ VRAM GGUF

    Qwen3.5-27B. Q4 ~16 GB. Also Qwen3.6-27B as a later dense refresh — same RAM class.

  3. Large Gemma 3 with vision — Mac-first in LM Mini (4-bit MLX), or a home GPU via Connect.

    27B ~15.5 GB (4-bit) Phone — Mac 24–36 GB GPU 24 GB+ VRAM MLX · GGUF

    mlx-community/gemma-3-27b-it-4bit in LM Mini (~15.5 GB). Needs ~24 GB unified memory.

  4. Mac-first Gemma 3 — premium quality when you have unified memory to spare.

    12B ~7.1 GB Phone — Mac 24 GB+ GPU 16 GB+ VRAM MLX · GGUF

    google/gemma-3-12b-it. MLX 4-bit in LM Mini on Mac.

  5. Gemma 4 12B Instruct

    LM Studio / Ollama

    2026 Gemma 4 12B — text, image, and audio. Mac or GPU; too big for phones.

    12B ~7–8 GB (Q4/QAT) Phone — Mac 24 GB+ GPU 16 GB+ VRAM GGUF

    google/gemma-4-12B-it. Official QAT GGUF. 256K context. Host in LM Studio then chat from LM Mini.

  6. Qwen 3 14B

    In LM Mini

    Desktop-quality Qwen 3 in LM Mini on Mac — 4-bit MLX on 16 GB+ unified memory.

    14B ~8.2 GB (4-bit) Phone — Mac 16–32 GB GPU 12–16 GB VRAM MLX · GGUF

    mlx-community/Qwen3-14B-4bit in LM Mini (~8.2 GB). GGUF Q4 similar. Phone: Connect only.

  7. Qwen 2.5 14B Instruct

    LM Studio / Ollama

    Serious desktop quality — best on a home GPU or a high-RAM Mac.

    14B ~8–9 GB (Q4) Phone — Mac 32 GB+ GPU 12–16 GB VRAM GGUF · MLX

    Qwen2.5-14B-Instruct Q4/Q5; LM Studio / Ollama.

  8. DeepSeek R1 Distill 14B

    LM Studio / Ollama

    Heavy reasoning model for research-style answers on GPU or big Macs.

    14B ~8–9 GB (Q4) Phone — Mac 32 GB+ GPU 16 GB+ VRAM GGUF

    R1 14B distill; long context CoT; prefer desktop.

Home GPU / remote via Connect

Desktop-class models. Run in LM Studio or Ollama on a PC, then chat from your phone with LM Mini Connect.

  1. Mistral Small 3.2 24B

    LM Studio / Ollama

    2025 Mistral Small — rich writing and analysis. GPU or 48 GB+ Mac.

    24B ~13 GB (Q4) Phone — Mac 48 GB+ GPU 16–24 GB VRAM GGUF

    Mistral-Small-3.2-24B-Instruct-2506. Q4 ~13 GB.

  2. Qwen 2.5 32B Instruct

    LM Studio / Ollama

    Near-frontier open weights for a well-equipped GPU PC.

    32B ~18 GB (Q4) Phone — Mac 64 GB+ GPU 24 GB+ VRAM GGUF

    Qwen2.5-32B-Instruct; Q4 ~18 GB; LM Studio remote.

  3. Qwen 3 32B

    In LM Mini

    Top in-app Qwen 3 on 32 GB+ Macs (4-bit MLX). Phones should Connect, not download.

    32B ~18 GB (4-bit) Phone — Mac 32–64 GB GPU 24 GB+ VRAM MLX · GGUF

    mlx-community/Qwen3-32B-4bit in LM Mini (~18 GB). minRam 32 GB.

  4. Llama 4 Scout

    LM Studio / Ollama

    Meta’s Llama 4 Scout MoE (17B×16E). Home GPU / big Mac — then chat from the phone via Connect.

    109B ~40 GB+ (Q4 MoE) Phone — Mac — (use remote) GPU 48 GB+ VRAM GGUF

    Llama-4-Scout-17B-16E-Instruct. Mixture-of-experts; Q4 still tens of GB.

  5. Llama 4 Maverick

    LM Studio / Ollama

    Meta’s larger Llama 4 MoE (17B×128E). Multi-GPU / 80 GB class — Connect from the phone; do not download to the handset.

    400B ~80 GB+ (Q4 MoE) Phone — Mac — (use remote) GPU 80 GB+ VRAM GGUF

    Llama-4-Maverick-17B-128E-Instruct. Far heavier than Scout. FP8 and GGUF community quants.

  6. Gemma 4 26B-A4B Instruct

    LM Studio / Ollama

    MoE Gemma 4: 26B total, ~4B active. Big-model taste at 8B-class compute — GPU or 32 GB+ Mac, then Connect.

    26B ~8–12 GB (Q4 MoE) Phone — Mac 32 GB+ GPU 16 GB+ VRAM GGUF

    google/gemma-4-26B-A4B-it. Official QAT GGUF. Efficiency king of the Gemma 4 stack.

  7. Gemma 4 31B Instruct

    LM Studio / Ollama

    Dense Gemma 4 flagship — reasoning, vision, coding. Home GPU / 48 GB Mac, chat from the phone via Connect.

    31B ~18 GB (Q4/QAT) Phone — Mac 48 GB+ GPU 24 GB+ VRAM GGUF

    google/gemma-4-31B-it. Official QAT GGUF. 256K context. Apache 2.0.

  8. Qwen 3.5 35B-A3B

    LM Studio / Ollama

    2026 Qwen MoE: 35B total, ~3B active. Strong Mac/GPU pick — not a phone download.

    35B ~12–16 GB (Q4 MoE) Phone — Mac 32 GB+ GPU 16–24 GB VRAM GGUF

    Qwen3.5-35B-A3B. Lighter than dense 27B at similar quality. Qwen3.6-35B-A3B is the later FP8/MoE refresh in the same RAM band.

  9. Llama 3.3 70B Instruct

    LM Studio / Ollama

    Top-tier Meta 70B — run at home on a strong GPU, chat from your phone via Connect.

    70B ~40 GB (Q4) Phone — Mac — (use remote) GPU 48 GB+ VRAM GGUF

    Llama-3.3-70B-Instruct; Q4 ~40 GB; multi-GPU friendly.

  10. Qwen 2.5 72B Instruct

    LM Studio / Ollama

    Huge Qwen for maximum quality on a serious home or office GPU rig.

    72B ~40+ GB (Q4) Phone — Mac — (use remote) GPU 48 GB+ VRAM GGUF

    Qwen2.5-72B-Instruct; Q4 ~40+ GB; Ollama / LM Studio.

  11. DeepSeek R1 Distill 70B

    LM Studio / Ollama

    Research-grade reasoning — pair with Connect so your phone stays light.

    70B ~40+ GB (Q4) Phone — Mac — (use remote) GPU 48 GB+ VRAM GGUF

    70B R1 distill; expects large VRAM or multi-GPU.

  12. Mixtral 8x7B Instruct

    LM Studio / Ollama

    Mixture-of-experts — high quality without always loading a dense 70B.

    46.7B ~26 GB (Q4) / less active Phone — Mac 64 GB+ GPU 24 GB+ VRAM GGUF

    Sparse MoE; ~12–26 GB depending on quant; LM Studio classic.