OfflineLingo runs whisper.cpp and llama.cpp on a phone with no network permission at all. Gemmory runs Gemma 4 through LiteRT-LM with every note and answer kept in a local Room database. The lessons overlap almost completely.

Explore this article visually3 figures

01

Separate waiting from working

02

The budget comes before the model

Figure 01Architecture

What actually occupies the phone’s memory

Select an element to explore its role.
Read every explanation
Application
The interface, audio capture and persistence need headroom while inference is running. Model weights do not get the entire process budget.
Model weights
Quantisation changes weight storage, but file size alone does not describe resident memory or temporary loading allocations.
Working memory
Context length and runtime choices affect working memory. Profile loading and generation rather than inferring the peak from model download sizes.
Operating margin
The operating system can reclaim memory under pressure. A foreground service is not a guarantee that models will remain resident.
01 / 04
Application

The interface, audio capture and persistence need headroom while inference is running. Model weights do not get the entire process budget.

Conceptual memory layers, not a measured allocation chart. Peak usage depends on the device and runtime.

The first question is what the target phone can sustain. Android does not promise an app a fixed share of physical RAM: heap limits, native allocations, other processes and system memory pressure all matter. Model downloads also compete with the user’s remaining storage. Set a budget on named devices, then measure the complete app against it.

ConstraintQuestion I wrote downWhat it ruled out
Resident RAMCan both models stay loaded between phrases?whisper small (≈ 466 MB f16) plus a 3B language model
First responseHow long until the person sees text?Any pipeline that waits for the full translation before rendering
StorageWhat is the install-plus-models footprint?Shipping several language pairs as separate models
Battery / thermalCan it run for a 20-minute conversation?Running the language model at full context on every phrase

For OfflineLingo the answer became whisper.cpp with a quantised base model (about 60 MB) and llama.cpp with a 4-bit Qwen 2.5 instruct model of roughly 1 GB. Neither is the best model available. Together they are the largest pair that leaves enough headroom for the app not to be killed while the user is mid-sentence.

03

Loading is the feature nobody demos

A local model has a lifecycle that a cloud API hides: it must be downloaded, verified, mapped into memory and kept warm. Every one of those steps fails in a way the user can see.

Gemmory verifies the model file's size and SHA-256 before loading it, because a partial download that loads and then crashes on the first token is far worse than a clear "model incomplete, resume download" state. The check costs a couple of seconds once and removes an entire class of impossible-to-reproduce crashes.

kotlinVerify before load. A model that half-loads is a crash the user cannot explain.
suspend fun ensureModel(spec: ModelSpec): ModelState {
    val file = File(context.filesDir, spec.fileName)
    if (!file.exists() || file.length() != spec.sizeBytes) return ModelState.Missing
    val digest = withContext(Dispatchers.IO) { sha256(file) }
    if (digest != spec.sha256) { file.delete(); return ModelState.Corrupt }
    return ModelState.Ready(file)
}

Keeping models warm matters just as much. Loading a 1 GB model takes several seconds even with mmap; doing that per phrase would make the app unusable. Both models are loaded once, held in a foreground service, and released only under memory pressure—at which point the UI says so instead of silently getting slower.

Diagram 02

A latency budget on one timeline

0s1s2s3s4s5s6sMic capture · 16 kHz2.5sstream + VADwhisper.cpp · base Q5_10.9s≈ 60 MBllama.cpp · Qwen 2.5 Q4_K_M1.4s≈ 1.0 GBRender + review0.8spersonfirst visible text (transcript)6 s budgetSequential budget: 5.6 s total · 3.1 s after capture, including review.Illustrative budget · measure peak memory, cold start and thermal pressure on the target device
Illustrative sequential budget: 5.6 seconds including capture and review, with translation complete at 4.8 seconds. Each stage contributes to the total; these are targets, not device measurements.

04

Degrade in one direction only

Figure 03Data flow

The model has a lifecycle before its first token

Select an element to explore its role.
Read every explanation
Acquire
Make setup explicit. A device without network permission needs a local import path; an online installer needs a resumable download.
Verify
Compare the completed file with a trusted manifest before loading. A mismatch should expose a recoverable error, not a generation crash.
Load
Keep the interface responsive while mapping and preparing the model. Measure cold start separately from warm inference.
Release
Explain when memory pressure forces unloading. The next request returns through loading instead of appearing mysteriously slower.
01 / 04
Acquire

Make setup explicit. A device without network permission needs a local import path; an online installer needs a resumable download.

A recoverable loading sequence. Corrupt or incomplete files return to acquisition, not to inference.

When the budget is exceeded, the app has to give something up, and it has to give up the same thing every time. On OfflineLingo the order is fixed: shorten the audio window first (12 s → 8 s), then drop whisper from base to tiny, then refuse new recordings until memory is back. The order never goes the other way and never involves a network—there is no network. The user sees a small mode indicator change; they never see a translation that silently came from a smaller model without a hint that it did.

Gemmory's version of the same rule is cancellable generation. A long answer streaming token by token can be stopped at any point, and the partial answer is kept as a note draft rather than discarded. The person is never waiting on something they cannot interrupt.

05

Privacy you can point at

A privacy policy is a claim. A missing permission is a fact. OfflineLingo's manifest does not declare android.permission.INTERNET, which means the operating system will refuse any socket the app tries to open. The strongest privacy statement in the project is one line that is not there.

Airplane mode is a useful functional test: it demonstrates that the installed models can support the interaction without connectivity. It does not prove that an app never transmits data when a network returns. Inspect the merged manifest, backup settings, exported components and any delegated actions separately. Also test a fresh installation with models imported locally; a cached-model demo does not cover setup.

06

What offline taught me about online products

The habits forced by having no network are the same habits that make connected products calm: state is explicit, every wait has a visible cause, every failure has a fixed next step, and nothing important is lost when a call does not return. I now write the degraded path for cloud-backed features the same way I did for OfflineLingo, and the interfaces are better for it even when the network is fine.

07

Measure a session, not a model file

File size is not resident memory. Measure peak process memory during loading and generation, including the KV cache, audio buffers and temporary allocations. Record the device, OS, runtime revision, model checksum, quantisation and context limit. Compare cold start, warm response and a sustained conversation; report the slow tail as well as the median. Choose the smallest model that meets the task-quality target with enough headroom for those conditions.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read