OfflineLingo runs whisper.cpp and llama.cpp on a phone with no network permission at all. Gemmory runs Gemma 4 through LiteRT-LM with every note and answer kept in a local Room database. The lessons overlap almost completely.
Explore this article visually3 figures
01
Separate waiting from working
02
The budget comes before the model
Figure 01Architecture
What actually occupies the phone’s memory
Read every explanation
- Application
- The interface, audio capture and persistence need headroom while inference is running. Model weights do not get the entire process budget.
- Model weights
- Quantisation changes weight storage, but file size alone does not describe resident memory or temporary loading allocations.
- Working memory
- Context length and runtime choices affect working memory. Profile loading and generation rather than inferring the peak from model download sizes.
- Operating margin
- The operating system can reclaim memory under pressure. A foreground service is not a guarantee that models will remain resident.
The interface, audio capture and persistence need headroom while inference is running. Model weights do not get the entire process budget.
The first question is what the target phone can sustain. Android does not promise an app a fixed share of physical RAM: heap limits, native allocations, other processes and system memory pressure all matter. Model downloads also compete with the user’s remaining storage. Set a budget on named devices, then measure the complete app against it.
| Constraint | Question I wrote down | What it ruled out |
|---|---|---|
| Resident RAM | Can both models stay loaded between phrases? | whisper small (≈ 466 MB f16) plus a 3B language model |
| First response | How long until the person sees text? | Any pipeline that waits for the full translation before rendering |
| Storage | What is the install-plus-models footprint? | Shipping several language pairs as separate models |
| Battery / thermal | Can it run for a 20-minute conversation? | Running the language model at full context on every phrase |
For OfflineLingo the answer became whisper.cpp with a quantised base model (about 60 MB) and llama.cpp with a 4-bit Qwen 2.5 instruct model of roughly 1 GB. Neither is the best model available. Together they are the largest pair that leaves enough headroom for the app not to be killed while the user is mid-sentence.
03
Loading is the feature nobody demos
A local model has a lifecycle that a cloud API hides: it must be downloaded, verified, mapped into memory and kept warm. Every one of those steps fails in a way the user can see.
Gemmory verifies the model file's size and SHA-256 before loading it, because a partial download that loads and then crashes on the first token is far worse than a clear "model incomplete, resume download" state. The check costs a couple of seconds once and removes an entire class of impossible-to-reproduce crashes.
suspend fun ensureModel(spec: ModelSpec): ModelState {
val file = File(context.filesDir, spec.fileName)
if (!file.exists() || file.length() != spec.sizeBytes) return ModelState.Missing
val digest = withContext(Dispatchers.IO) { sha256(file) }
if (digest != spec.sha256) { file.delete(); return ModelState.Corrupt }
return ModelState.Ready(file)
}Keeping models warm matters just as much. Loading a 1 GB model takes several seconds even with mmap; doing that per phrase would make the app unusable. Both models are loaded once, held in a foreground service, and released only under memory pressure—at which point the UI says so instead of silently getting slower.
Diagram 02
A latency budget on one timeline
04
Degrade in one direction only
Figure 03Data flow
The model has a lifecycle before its first token
Read every explanation
- Acquire
- Make setup explicit. A device without network permission needs a local import path; an online installer needs a resumable download.
- Verify
- Compare the completed file with a trusted manifest before loading. A mismatch should expose a recoverable error, not a generation crash.
- Load
- Keep the interface responsive while mapping and preparing the model. Measure cold start separately from warm inference.
- Release
- Explain when memory pressure forces unloading. The next request returns through loading instead of appearing mysteriously slower.
Make setup explicit. A device without network permission needs a local import path; an online installer needs a resumable download.
When the budget is exceeded, the app has to give something up, and it has to give up the same thing every time. On OfflineLingo the order is fixed: shorten the audio window first (12 s → 8 s), then drop whisper from base to tiny, then refuse new recordings until memory is back. The order never goes the other way and never involves a network—there is no network. The user sees a small mode indicator change; they never see a translation that silently came from a smaller model without a hint that it did.
Gemmory's version of the same rule is cancellable generation. A long answer streaming token by token can be stopped at any point, and the partial answer is kept as a note draft rather than discarded. The person is never waiting on something they cannot interrupt.
05
Privacy you can point at
A privacy policy is a claim. A missing permission is a fact. OfflineLingo's manifest does not declare android.permission.INTERNET, which means the operating system will refuse any socket the app tries to open. The strongest privacy statement in the project is one line that is not there.
Airplane mode is a useful functional test: it demonstrates that the installed models can support the interaction without connectivity. It does not prove that an app never transmits data when a network returns. Inspect the merged manifest, backup settings, exported components and any delegated actions separately. Also test a fresh installation with models imported locally; a cached-model demo does not cover setup.
06
What offline taught me about online products
The habits forced by having no network are the same habits that make connected products calm: state is explicit, every wait has a visible cause, every failure has a fixed next step, and nothing important is lost when a call does not return. I now write the degraded path for cloud-backed features the same way I did for OfflineLingo, and the interfaces are better for it even when the network is fine.
07
Measure a session, not a model file
File size is not resident memory. Measure peak process memory during loading and generation, including the KV cache, audio buffers and temporary allocations. Record the device, OS, runtime revision, model checksum, quantisation and context limit. Compare cold start, warm response and a sustained conversation; report the slow tail as well as the median. Choose the smallest model that meets the task-quality target with enough headroom for those conditions.
More articles
09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read