Speech recognition models give you several uncertainty signals for free. The design work is deciding which of them should change what a person sees, and making the correction cheaper than the doubt.

Explore this article visually3 figures

01

Test the action, not just the highlight

02

A 0.61 is not a 0.61

Figure 01Comparison

A token score and a segment diagnostic are different signals

Select an element to explore its role.
Read every explanation
Local ambiguity
One word may contain several tokens. A low token score can guide inspection, but does not provide a word-level confidence or alternative list on its own.
Segment quality
A segment-level signal concerns the whole decoded stretch. Noise or overlapping voices may make sentence replay more useful than highlighting one token.
Speech presence
Combine the no-speech diagnostic with other evidence. Do not silently discard a real recording because one threshold was crossed.
Repetition
Repetition can signal a decoding loop, but genuine speech can repeat too. Preserve context and give the person a recovery action.
01 / 04
Local ambiguity

One word may contain several tokens. A low token score can guide inspection, but does not provide a word-level confidence or alternative list on its own.

Diagnostic categories, not calibrated probabilities of correctness. Check which fields the runtime exposes.

The example that fixed this for me came from testing OfflineLingo with medical phrases. The transcript read “give him two tablets” and whisper's per-token probability on “two” was about 0.6, with “to” as the runner-up. A few words later an “uh” had the same probability. Showing both in the same shade of yellow would have been technically honest and practically useless: one is noise, the other could change a dose.

So the design question is not “how do we show confidence” but “what is the cost of this specific token being wrong, and what is the cheapest way for the person to check it”. Confidence is one input to that decision. Token type is the other, and it matters more.

03

Four signals worth distinguishing

The Whisper ecosystem offers several diagnostic signals, but bindings do not expose identical fields. The Python implementation uses the segment diagnostics below; in whisper.cpp, check the pinned API and compute missing diagnostics explicitly. Alternatives also require decoder support; a token score alone does not provide an n-best word list.

SignalWhat it usually meansReasonable response
Token probability (per subword)This word is ambiguous or the decoded token has limited supportUnderline; offer the n-best alternatives on tap
avg_logprob (per segment)The whole segment is shaky—noise, accent, crosstalkMark the sentence, offer replay of the segment
no_speech_probThe model doubts there was speech at allCombine with segment confidence; offer re-recording
compression_ratioRepetitive output—the classic hallucination loopFlag repetition; check audio before discarding

Silence and repetition are warning signals, not proof that a transcript is false. Genuine speech can repeat, and noise can confuse speech detection. Combine diagnostics, preserve the audio for review and offer re-recording when the segment cannot be trusted. Do not silently delete speech on the strength of one threshold.

Diagram 02

Consequence × confidence, not confidence alone

Consequence if wrong ↓Token probability (whisper) →< 0.550.55 ≤ p < 0.80≥ 0.80Filler wordshowshowshowName, placeunderline + alternativesunderlineshowNumber, dose, negationblock + replayask to confirmunderline“Give him two tablets” — p(two) = 0.61 → same score as an “uh”, opposite action.Illustrative thresholds to calibrate; the score does not guarantee correctness.
The same probability lands in different cells depending on what kind of token it is. The row—what happens if this word is wrong—decides the interface response; the column only tunes it.

04

The consequence axis

Classifying tokens by consequence sounds like it needs a model. It mostly needs a regular expression and a short list. Numbers, negations (not, no, never, ne … pas), units, and capitalised tokens that are not sentence-initial are a starting heuristic, not a complete account of meaning. Written-out numbers and phrases need additional rules, and unclassified words may still be consequential.

kotlinCheap consequence classification. The model gives probability; this gives the row.
fun consequence(token: Token, index: Int): Consequence = when {
    token.text.any { it.isDigit() }                     -> Consequence.HIGH   // doses, counts, times
    token.text.lowercase() in NEGATIONS                  -> Consequence.HIGH   // "not allergic" vs "allergic"
    token.text.lowercase() in UNITS                      -> Consequence.HIGH   // mg, ml, km
    index > 0 && token.text.firstOrNull()?.isUpperCase() == true        -> Consequence.MEDIUM // names, places
    else                                                 -> Consequence.LOW
}

With that in place, the interface needs exactly four behaviours: show, underline, ask to confirm, block-and-replay. A low-consequence token never triggers more than an underline, whatever its score. A high-consequence token below the threshold blocks the sentence from being marked as final until the person has replayed it or edited it. Thresholds must be validated on representative recordings; the example policy does not establish a safe operating point.

05

Correction has to cost less than distrust

Figure 03Data flow

Make the correction shorter than the doubt

Select an element to explore its role.
Read every explanation
Notice
Mark a consequential ambiguity where it changes the next action. Avoid displaying a percentage that suggests unverified calibration.
Listen
Replay enough context to understand the phrase. Offer alternatives only when the decoder actually supplies them.
Edit
Edit the transcript without recapturing the entire phrase. Keep the original audio available to resolve later disagreements.
Review
Translate the corrected transcript and review the output. A corrected source word does not prove that the target-language sentence is faithful.
01 / 04
Notice

Mark a consequential ambiguity where it changes the next action. Avoid displaying a percentage that suggests unverified calibration.

A proposed interaction flow. Correcting the transcript still requires checking the new translation.

A cue that leads nowhere trains people to ignore cues. Every marked token in OfflineLingo is tappable: the tap shows the n-best alternatives whisper considered, and a second control replays 1.5 seconds of audio centred on the word. Choosing an alternative or retyping the word updates the translation; the rest of the sentence is not re-run. That last part matters—if fixing one word meant waiting for the whole pipeline again, nobody would fix words.

I also removed the percentage. An early build showed 61 % next to the word and testers spent time reasoning about the number. The underline plus alternatives conveyed the same doubt and led directly to the action. A number invites interpretation; an underline invites a tap.

06

What I would measure next

  • Edit rate on flagged versus unflagged high-consequence tokens. If people edit unflagged ones as often, the flags are miscalibrated.
  • Time from flag to correction. Above a few seconds, the correction path is too heavy.
  • Flag fatigue: how many flags per sentence before people stop tapping. My guess is three; I would rather know.
  • How often compression_ratio catches a hallucinated segment in real noise. If it is frequent, that check deserves a visible state of its own.

07

A score needs a calibration set

A token probability is conditional on the audio and previously decoded tokens; it is not a calibrated probability that a word is correct. A word may contain several tokens. Validate any aggregation and thresholds on labelled recordings from the intended languages and acoustic conditions. Measure missed consequential errors and unnecessary interruptions separately. A simple token classifier also misses written-out numbers, multiword negations and names without capitals: unknown cases need review, not an automatic low-risk label.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read