Model guide
Which model should I run?
A plain-language list of 51 local models. Read the first sentence if you are new; the grey line is for specs. Download smaller ones in LM Mini, or load bigger ones on a home PC with LM Studio / Ollama and connect remotely. Not sure about RAM? Find your phone.
Everyday phones
Small models that fit most iPhones and Androids. Start here if you are new to local AI.
-
Qwen 3 0.6B
In LM MiniTiny and quick — good for short replies when your phone is low on memory.
Qwen3-0.6B; Q4 GGUF & 4-bit MLX. Hugging Face downloads in the tens of millions.
-
Qwen 3.5 0.8B
LM Studio / Ollama2026 Qwen tiny — hybrid long-context brain in a phone-sized file. Best when you want the newest Qwen on 4–6 GB devices.
Qwen3.5-0.8B (Feb 2026). Gated DeltaNet hybrid; 262K native context. Community GGUF/MLX as engines catch up — import in LM Mini or host via Connect.
-
Gemma 3 1B Instruct
In LM MiniGoogle’s smallest Gemma 3 chat model. Great starter for everyday questions.
google/gemma-3-1b-it. Instruct-tuned; GGUF + MLX.
-
Gemma 4 E2B Instruct
LM Studio / OllamaGoogle’s 2026 edge Gemma — ~2.3B effective, text + image + audio. The new tiny phone default if your engine supports Gemma 4.
google/gemma-4-E2B-it (Apr 2026). Official QAT GGUF. Apache 2.0. ~5.1B with embeddings; plan ~2–3 GB live.
-
Llama 3.2 1B Instruct
In LM MiniMeta’s pocket Llama — fast, tool-friendly, and solid for light chat.
Llama-3.2-1B-Instruct; tool calling. GGUF + MLX.
-
SmolLM2 1.7B Instruct
LM Studio / OllamaHugging Face’s small all-rounder. Surprisingly capable for its size.
SmolLM2-1.7B-Instruct Q4_K_M via LM Studio / custom HF import.
-
DeepSeek R1 Distill 1.5B
In LM MiniA mini “thinking” model — stronger at math and step-by-step reasoning.
R1 distill of Qwen 2.5 1.5B; reasoning traces. GGUF + MLX.
-
Qwen 3 1.7B
In LM MiniBest everyday phone pick for most people — chat, tools, and general help.
Qwen3-1.7B; primary free-slot family in LM Mini. GGUF + MLX.
-
Qwen 3.5 2B
LM Studio / OllamaNewest small Qwen that still fits everyday phones — sharper than 1.7B-class, still downloadable on cellular.
Qwen3.5-2B. Hybrid attention + native multimodal in the 3.5 family. Prefer Q4 GGUF once your llama.cpp/MLX build lists Qwen3.5.
-
Qwen 2.5 1.5B Instruct
LM Studio / OllamaPrevious-gen Qwen that’s still a reliable tiny assistant.
Qwen2.5-1.5B-Instruct; still huge Hugging Face download volume.
-
Gemma 3n E2B
LM Studio / OllamaGoogle’s on-device Gemma 3n (effective 2B). Built for phones, including audio-aware builds.
google/gemma-3n-E2B-it. LiteRT / GGUF community quants. Newer than Gemma 3 1B.
Pro phones & tablets
More capable answers — needs roughly 6 GB+ free RAM on the device.
-
Llama 3.2 3B Instruct
In LM MiniNoticeably smarter than 1B — good for writing help and longer chats on Pro phones.
Llama-3.2-3B-Instruct; tool calling. GGUF + MLX.
-
SmolLM3 3B
LM Studio / OllamaHugging Face’s 2025 small model — a step up from SmolLM2, still phone-sized.
HuggingFaceTB/SmolLM3-3B. GGUF via bartowski / community.
-
Gemma 3 4B Instruct
In LM MiniGoogle’s mid-size Gemma 3 — sharper answers; some builds can look at images.
google/gemma-3-4b-it; multimodal on supported builds. GGUF + MLX.
-
Qwen 3.5 4B
LM Studio / Ollama2026’s 4B daily driver — long context and stronger tools than Qwen 3 4B, if your phone has ~6 GB free.
Qwen3.5-4B; millions of HF downloads. Hybrid Gated DeltaNet. Import a Q4 when the engine supports it; Qwen 3 4B Instruct 2507 remains the in-app pick today.
-
Gemma 4 E4B Instruct
LM Studio / OllamaGoogle’s on-device Gemma 4 (~4.5B effective) with image and audio. The 2026 upgrade from Gemma 3 4B / 3n E4B.
google/gemma-4-E4B-it. Official QAT GGUF. ~8B with embeddings; budget ~4 GB live. Apache 2.0.
-
Phi-4 Mini Instruct
In LM MiniStrong reasoning in a phone-friendly package — great for explanations and code-ish help.
microsoft/Phi-4-mini-instruct (~3.8B). GGUF + MLX.
-
Qwen 3 4B Instruct 2507
In LM MiniBalanced daily driver for Pro phones and Macs — quality without a desktop GPU.
July 2025 Instruct refresh; millions of HF downloads. GGUF + MLX.
-
Qwen 2.5 3B Instruct
LM Studio / OllamaCompact Qwen for chat and light coding when 4B feels heavy.
Qwen2.5-3B-Instruct Q4_K_M; LM Studio / Ollama.
-
Gemma 3n E4B
LM Studio / OllamaGemma 3n effective-4B — Google’s newer on-device stack, including multimodal 3n builds.
google/gemma-3n-E4B-it. Community GGUF (bartowski). Heavier than E2B.
-
Nemotron 3 Nano 4B
LM Studio / OllamaNVIDIA’s 2026 tiny Nemotron — dense 4B-class, aimed at local agents.
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Use a GGUF quant on phone; BF16 is a desktop download.
Mac (M-series) & strong phones
8B-class models. Comfortable on 16 GB+ Macs; some high-RAM phones can run them too.
-
Mistral 7B Instruct
LM Studio / OllamaClassic open model — clear writing and solid general knowledge on Mac or GPU.
Mistral-7B-Instruct v0.3; ubiquitous GGUF quants.
-
Llama 3.1 8B Instruct
LM Studio / OllamaMeta’s workhorse 8B — excellent general assistant on Mac or a home GPU.
Llama-3.1-8B-Instruct; tool calling; LM Studio / Ollama staple.
-
Qwen 3 8B
In LM MiniFrontier feel on-device for high-RAM phones and M-series Macs.
Qwen3-8B; tool calling. GGUF + MLX in LM Mini.
-
Qwen 3.5 9B
LM Studio / OllamaFrontier feel on a 16 GB phone or any M-series Mac — the 2026 8B-class upgrade.
Qwen3.5-9B (most-downloaded 3.5 size on HF). Not in the LM Mini in-app list yet; LM Studio / Ollama / custom import + Connect.
-
Qwen 2.5 7B Instruct
LM Studio / OllamaPopular all-rounder for coding help, summaries, and chat on desktop or Mac.
Qwen2.5-7B-Instruct — still one of the most downloaded instruct models on HF.
-
DeepSeek R1 Distill 8B
LM Studio / OllamaReasoning-focused — walks through hard problems before answering.
R1 distill on Llama 8B; longer CoT traces.
-
OLMo 3 7B Instruct
LM Studio / OllamaAllen AI’s fully open 7B instruct model — a transparent alternative to Llama/Qwen.
allenai/Olmo-3-7B-Instruct (2025). Community GGUF.
-
Nemotron Nano 9B v2
LM Studio / OllamaNVIDIA 9B local model — stronger than 8B class if you have the RAM.
nvidia/NVIDIA-Nemotron-Nano-9B-v2. Prefer Q4 GGUF on Mac/GPU.
-
Gemma 2 9B Instruct
LM Studio / OllamaGoogle’s 9B chat model — polished answers when you have Mac RAM or a GPU.
Gemma-2-9B-IT; still excellent instruction following.
-
Hermes 3 8B
LM Studio / OllamaCommunity favorite for creative writing and agent-style tool use.
NousResearch Hermes-3-Llama-3.1-8B; function calling.
High-memory Mac
12B–14B class. Best on 24–32 GB+ unified memory, or remote from a GPU PC.
-
Qwen 3 30B-A3B Instruct
LM Studio / OllamaMixture-of-experts: 30B total, ~3B active. Big-model taste on a strong Mac or GPU.
Qwen3-30B-A3B-Instruct-2507 MoE. Q4 is lighter than dense 30B; still not a phone model.
-
Qwen 3.5 27B
LM Studio / OllamaDense 2026 Qwen for a high-RAM Mac or a 24 GB GPU — then use the phone as the remote keyboard.
Qwen3.5-27B. Q4 ~16 GB. Also Qwen3.6-27B as a later dense refresh — same RAM class.
-
Gemma 3 27B Instruct
In LM MiniLarge Gemma 3 with vision — Mac-first in LM Mini (4-bit MLX), or a home GPU via Connect.
mlx-community/gemma-3-27b-it-4bit in LM Mini (~15.5 GB). Needs ~24 GB unified memory.
-
Gemma 3 12B Instruct
In LM MiniMac-first Gemma 3 — premium quality when you have unified memory to spare.
google/gemma-3-12b-it. MLX 4-bit in LM Mini on Mac.
-
Gemma 4 12B Instruct
LM Studio / Ollama2026 Gemma 4 12B — text, image, and audio. Mac or GPU; too big for phones.
google/gemma-4-12B-it. Official QAT GGUF. 256K context. Host in LM Studio then chat from LM Mini.
-
Qwen 3 14B
In LM MiniDesktop-quality Qwen 3 in LM Mini on Mac — 4-bit MLX on 16 GB+ unified memory.
mlx-community/Qwen3-14B-4bit in LM Mini (~8.2 GB). GGUF Q4 similar. Phone: Connect only.
-
Qwen 2.5 14B Instruct
LM Studio / OllamaSerious desktop quality — best on a home GPU or a high-RAM Mac.
Qwen2.5-14B-Instruct Q4/Q5; LM Studio / Ollama.
-
DeepSeek R1 Distill 14B
LM Studio / OllamaHeavy reasoning model for research-style answers on GPU or big Macs.
R1 14B distill; long context CoT; prefer desktop.
Home GPU / remote via Connect
Desktop-class models. Run in LM Studio or Ollama on a PC, then chat from your phone with LM Mini Connect.
-
Mistral Small 3.2 24B
LM Studio / Ollama2025 Mistral Small — rich writing and analysis. GPU or 48 GB+ Mac.
Mistral-Small-3.2-24B-Instruct-2506. Q4 ~13 GB.
-
Qwen 2.5 32B Instruct
LM Studio / OllamaNear-frontier open weights for a well-equipped GPU PC.
Qwen2.5-32B-Instruct; Q4 ~18 GB; LM Studio remote.
-
Qwen 3 32B
In LM MiniTop in-app Qwen 3 on 32 GB+ Macs (4-bit MLX). Phones should Connect, not download.
mlx-community/Qwen3-32B-4bit in LM Mini (~18 GB). minRam 32 GB.
-
Llama 4 Scout
LM Studio / OllamaMeta’s Llama 4 Scout MoE (17B×16E). Home GPU / big Mac — then chat from the phone via Connect.
Llama-4-Scout-17B-16E-Instruct. Mixture-of-experts; Q4 still tens of GB.
-
Llama 4 Maverick
LM Studio / OllamaMeta’s larger Llama 4 MoE (17B×128E). Multi-GPU / 80 GB class — Connect from the phone; do not download to the handset.
Llama-4-Maverick-17B-128E-Instruct. Far heavier than Scout. FP8 and GGUF community quants.
-
Gemma 4 26B-A4B Instruct
LM Studio / OllamaMoE Gemma 4: 26B total, ~4B active. Big-model taste at 8B-class compute — GPU or 32 GB+ Mac, then Connect.
google/gemma-4-26B-A4B-it. Official QAT GGUF. Efficiency king of the Gemma 4 stack.
-
Gemma 4 31B Instruct
LM Studio / OllamaDense Gemma 4 flagship — reasoning, vision, coding. Home GPU / 48 GB Mac, chat from the phone via Connect.
google/gemma-4-31B-it. Official QAT GGUF. 256K context. Apache 2.0.
-
Qwen 3.5 35B-A3B
LM Studio / Ollama2026 Qwen MoE: 35B total, ~3B active. Strong Mac/GPU pick — not a phone download.
Qwen3.5-35B-A3B. Lighter than dense 27B at similar quality. Qwen3.6-35B-A3B is the later FP8/MoE refresh in the same RAM band.
-
Llama 3.3 70B Instruct
LM Studio / OllamaTop-tier Meta 70B — run at home on a strong GPU, chat from your phone via Connect.
Llama-3.3-70B-Instruct; Q4 ~40 GB; multi-GPU friendly.
-
Qwen 2.5 72B Instruct
LM Studio / OllamaHuge Qwen for maximum quality on a serious home or office GPU rig.
Qwen2.5-72B-Instruct; Q4 ~40+ GB; Ollama / LM Studio.
-
DeepSeek R1 Distill 70B
LM Studio / OllamaResearch-grade reasoning — pair with Connect so your phone stays light.
70B R1 distill; expects large VRAM or multi-GPU.
-
Mixtral 8x7B Instruct
LM Studio / OllamaMixture-of-experts — high quality without always loading a dense 70B.
Sparse MoE; ~12–26 GB depending on quant; LM Studio classic.