Aug 28, 2026 · 4 min read

Which Models Run Well on Your Phone?

Which Models Run Well on Your Phone?

If you only remember one sentence from this post, make it this one: on a phone, the “best” model is the one that answers before you get bored and still leaves memory for the rest of your life (messages, maps, the app that is definitely not a game).

The short version of how this works

On-device in LM Mini means the model file lives on your iPhone, iPad, Android, or Mac and runs there. Formats people hit most often are things like GGUF and MLX, depending on platform. You download a model, load it, and chat offline if you want.

That sounds magical until RAM shows up with a clipboard and a disappointed face.

Yes, the RAM rant

Phones do not have desktop RAM culture. Your laptop might shrug at a 12GB model. Your phone is also running the OS, your chat UI, maybe voice, maybe a Persona with memories, and whatever notifications are vibrating in your pocket.

So when a model is “only” a few gigabytes on disk, the runtime can still want a serious chunk of memory. If you push it, you get heat, throttling, slow first tokens, or the app politely falling over. Not because LM Mini hates you. Because physics and memory budgets are rude.

I have watched people download the biggest file they can find, then message support like the phone betrayed them. The phone did its best. The download button did not come with a warning label written in enough sarcasm.

What tends to feel good on-device

Rules of thumb, not commandments:

  • Start small. A compact model that streams quickly will teach you more about the app than a giant one that thinks for twenty seconds and then sighs.
  • Quantization matters. A well-quantized smaller model often beats a “bigger on paper” model that your device can barely hold.
  • Match the job. Quick rewrites, summaries, brainstorming, and tutoring can work great on-device. Deep multi-file coding and long research dumps are usually happier on a desktop GPU through LM Studio or Ollama.
  • Leave headroom. If the model loads but everything else gets laggy, size down. A chat app that freezes while your music stutters is not “powerful.” It is annoying.

Exact “use this one model forever” lists go stale fast, so I am not going to pretend a blog post from today is a permanent leaderboard. Check the in-app model options / docs, try two sizes, and keep the one you actually enjoy using. The phone × model pages on the site are a useful starting map.

Phone vs home PC (the honest split)

Use on-device when you want:

  • privacy with fewer moving parts
  • offline / travel / couch mode
  • fast answers for everyday questions
  • a Persona that does not need a datacenter brain

Use LM Studio / Ollama on your PC (with LM Mini on Wi‑Fi or via Home / Connect) when you want:

  • larger models
  • longer context
  • heavier coding or analysis
  • less worry about phone thermals

That is the whole product idea in one paragraph: the phone can be the computer, or the remote control. You pick based on the day, not based on internet arguments.

How I would test models in 10 minutes

  1. Pick a small/medium on-device model and load it.
  2. Ask the same three prompts every time: a short rewrite, a simple explanation, and one slightly harder reasoning question.
  3. Notice speed, quality, and whether the phone gets weirdly warm.
  4. Size up once. If quality jumps and speed is still fine, great. If quality barely changes and the phone turns into a hand warmer, go back.
  5. For “serious mode,” connect to your home LM Studio or Ollama model and compare. Usually the desktop one wins on depth, and the phone one wins on convenience.

If you are using Personas, bind lighter models to everyday characters and heavier desktop models to the picky reviewer Persona. That combo is underrated.

Common mistakes

  • Downloading the largest model first “just to see”
  • Judging a model on one unlucky answer
  • Forgetting that chat history + voice + tools all cost something too
  • Expecting phone models to match a big desktop setup token-for-token
  • Ignoring that a slightly “dumber” model you actually use beats a genius model you avoid because it is slow

The bottom line

On-device AI is real and useful, but your phone is not a silent workstation with infinite memory. Start smaller than your ego wants, upgrade only when you feel the limit, and keep LM Studio or Ollama ready for the big jobs.

That is how LM Mini stays fun instead of becoming a heat-and-hope experiment.

Next up in the series: LM Mini + iOS Shortcuts, for the people who want local AI one tap away from the home screen. Until then, try one smaller model tonight and notice how much nicer the app feels when RAM is not screaming.

Get LM Mini: lmmini.com