This is a project note rather than an essay—what we built, the decisions I would defend, the ones I would revisit, and the state machine that ended up mattering more than either model.

Explore this article visually3 figures

01

Correct the source while translation is running

02

The brief

The scenario given to us was cross-border emergency response: a responder needs to exchange short, high-stakes phrases with someone who does not share a language, in a place where the network may be gone. Not degraded—gone. That single constraint removed every cloud translation API from the design in the first hour and set the actual problem: fit speech recognition and translation onto a phone, and make the result something a stressed person can trust or correct in seconds.

03

What we shipped

Figure 01Data flow

A spoken phrase crosses three representations

Select an element to explore its role.
Read every explanation
Audio
Capture the original phrase with enough timing context for replay. A recording is the evidence; the transcript is an interpretation of it.
Transcript
Inspect numbers, units, names and negations before translation. Written-out numbers require more than a digit-matching expression.
Translation
Compare meaning rather than just matching digits. Equivalent units and number formatting can change while the meaning is preserved—or hide a real error.
01 / 03
Audio

Capture the original phrase with enough timing context for replay. A recording is the evidence; the transcript is an interpretation of it.

Each boundary can change meaning. Preserve the source representation so the person can inspect the transformation.
  • A Kotlin Android app with a single press-and-hold record control; releasing, or 1.5 seconds of silence, ends the phrase.
  • whisper.cpp through JNI with a quantised base model for on-device speech recognition, streaming 16 kHz PCM from the microphone.
  • llama.cpp with a 4-bit Qwen 2.5 instruct model for translation, kept loaded in a foreground service between phrases.
  • Per-word confidence from whisper's token probabilities, rendered as underlines with tap-to-see-alternatives and a replay of the surrounding audio.
  • A large-type output screen designed to be turned toward the other person, with the language pair visible at all times.
  • No INTERNET permission in the manifest. The OS enforces the promise, not the code.

04

The prompt is a contract too

Small instruct models like to be helpful. The first translation prompt returned things like “Here is the translation: …” or a translation followed by a note about ambiguity. In a large-type screen turned toward a stranger, that framing is noise at best and confusing at worst. The fix was to treat the prompt output like a tool result: constrain it, then validate it.

textThe translation prompt after three iterations. Short, and enforced by stop tokens and a post-check.
<|im_start|>system
You translate {src} to {dst} for emergency responders.
Output only the translated sentence. No preamble, no notes, no quotes.
Keep numbers, units and names exactly as given.
<|im_end|>
<|im_start|>user
{sentence}
<|im_end|>
<|im_start|>assistant

Generation stops at the first newline or end-of-turn token. A post-check compares digits and units in the input and output; if a number is missing on one side, the translation is shown with the number underlined in the same way a low-confidence word would be, because from the user's point of view it is the same problem. The line “keep numbers … exactly as given” helped, but the check is what made it dependable.

Diagram 02

The app as a state machine

IdleRecording≤ 12 sTranscribingwhisper.cppTranslatingQwen 2.5ReviewconfidenceEdit textShown · large typepress & holdrelease / 1.5 s silencesegments okfirst token < 1 sp < 0.55 on key wordfix texttranslate againacceptre-recordnext phrasefailure / 8 s timeoutEditing returns to translation; review offers a retry or acceptance.
Seven states, with the timeouts and thresholds that move between them. Dashed transitions are the ones triggered by uncertainty or failure; solid ones are the happy path and the person's own actions.

05

The state machine mattered more than the models

Halfway through the second day the pipeline worked and the app still felt unusable, because it could get stuck. Recording with no end, a transcription that returned nothing, a translation that took eleven seconds because the phone had throttled—each of these left the person looking at a screen with no obvious next move.

Drawing the app as a state machine fixed that faster than any model change. Every state where a person is waiting got two exits: a success transition and a bounded failure transition (a timeout or a threshold) that leads somewhere with a clear action. Recording ends at 12 seconds no matter what. Transcription or translation that exceeds 8 seconds returns to idle with a “try a shorter phrase” message rather than a spinner. A low-confidence key word routes to review with the word marked, instead of to the output screen. The diagram above is the version we shipped; the useful part of it is that no state can end without a transition out.

06

Decisions I would defend, decisions I would revisit

Defend:

  • Removing the network permission rather than adding an offline mode. A mode can be toggled; a missing permission cannot.
  • Press-and-hold recording. Tap-to-start/tap-to-stop produced long recordings with background chatter that whisper transcribed enthusiastically.
  • Marking numbers by consequence rather than by confidence alone; this caught more real errors in testing than the probability threshold did.

Revisit:

  • The base whisper model. small was noticeably better on accented speech, but the memory margin made me nervous. On a device with 6 GB or more I would use it.
  • One translation direction per screen. Responders need both directions in one conversation; the swap control was one tap too many.
  • A general-purpose instruct model for translation. A dedicated small translation model, or a fine-tune on emergency phrases, would likely be both faster and more literal.

07

What I would test next

Figure 03Comparison

Two interruptions, two useful exits

Select an element to explore its role.
Read every explanation
Generation times out
End the wait with a clear reason. Preserve the recording and transcript, and offer a shorter phrase or retry rather than forcing the person to start without context.
Meaning is uncertain
Let the person replay and edit before finalising. Finishing quickly is not success when the meaning remains unresolved.
01 / 02
Generation times out

End the wait with a clear reason. Preserve the recording and transcript, and offer a shorter phrase or retry rather than forcing the person to start without context.

Recovery design for a prototype. These paths describe interface behaviour, not validated field performance.

The prototype was tested by us, in a quiet hall, in languages we speak. The next test is the one that counts: outdoor noise, an accent the model has not seen, a phrase containing a dosage and a negation, on a phone that has been in a pocket for an hour, held by someone who has not seen the app before. If the state machine holds and the underlines land on the right words in those conditions, the project is worth taking further. If not, the diagram tells us exactly which transition to fix.

08

Turn the next test into a protocol

Use the same phrase set across devices and record the original audio, expected meaning, corrected transcript and final translation. Include numbers written as words, negations, interruptions and silence. Ask bilingual reviewers to score meaning preservation without seeing the model configuration. Report errors separately from completion time and abandonment. Replaying a word confirms what was heard; it does not validate a translation. This remains a prototype evaluation, not evidence of suitability for emergency use.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read