Speech recognition models give you several uncertainty signals for free. The design work is deciding which of them should change what a person sees, and making the correction cheaper than the doubt.
Explore this article visually3 figures
01
Test the action, not just the highlight
02
A 0.61 is not a 0.61
Figure 01Comparison
A token score and a segment diagnostic are different signals
Read every explanation
- Local ambiguity
- One word may contain several tokens. A low token score can guide inspection, but does not provide a word-level confidence or alternative list on its own.
- Segment quality
- A segment-level signal concerns the whole decoded stretch. Noise or overlapping voices may make sentence replay more useful than highlighting one token.
- Speech presence
- Combine the no-speech diagnostic with other evidence. Do not silently discard a real recording because one threshold was crossed.
- Repetition
- Repetition can signal a decoding loop, but genuine speech can repeat too. Preserve context and give the person a recovery action.
One word may contain several tokens. A low token score can guide inspection, but does not provide a word-level confidence or alternative list on its own.
The example that fixed this for me came from testing OfflineLingo with medical phrases. The transcript read “give him two tablets” and whisper's per-token probability on “two” was about 0.6, with “to” as the runner-up. A few words later an “uh” had the same probability. Showing both in the same shade of yellow would have been technically honest and practically useless: one is noise, the other could change a dose.
So the design question is not “how do we show confidence” but “what is the cost of this specific token being wrong, and what is the cheapest way for the person to check it”. Confidence is one input to that decision. Token type is the other, and it matters more.
03
Four signals worth distinguishing
The Whisper ecosystem offers several diagnostic signals, but bindings do not expose identical fields. The Python implementation uses the segment diagnostics below; in whisper.cpp, check the pinned API and compute missing diagnostics explicitly. Alternatives also require decoder support; a token score alone does not provide an n-best word list.
| Signal | What it usually means | Reasonable response |
|---|---|---|
| Token probability (per subword) | This word is ambiguous or the decoded token has limited support | Underline; offer the n-best alternatives on tap |
avg_logprob (per segment) | The whole segment is shaky—noise, accent, crosstalk | Mark the sentence, offer replay of the segment |
no_speech_prob | The model doubts there was speech at all | Combine with segment confidence; offer re-recording |
compression_ratio | Repetitive output—the classic hallucination loop | Flag repetition; check audio before discarding |
Silence and repetition are warning signals, not proof that a transcript is false. Genuine speech can repeat, and noise can confuse speech detection. Combine diagnostics, preserve the audio for review and offer re-recording when the segment cannot be trusted. Do not silently delete speech on the strength of one threshold.
Diagram 02
Consequence × confidence, not confidence alone
04
The consequence axis
Classifying tokens by consequence sounds like it needs a model. It mostly needs a regular expression and a short list. Numbers, negations (not, no, never, ne … pas), units, and capitalised tokens that are not sentence-initial are a starting heuristic, not a complete account of meaning. Written-out numbers and phrases need additional rules, and unclassified words may still be consequential.
fun consequence(token: Token, index: Int): Consequence = when {
token.text.any { it.isDigit() } -> Consequence.HIGH // doses, counts, times
token.text.lowercase() in NEGATIONS -> Consequence.HIGH // "not allergic" vs "allergic"
token.text.lowercase() in UNITS -> Consequence.HIGH // mg, ml, km
index > 0 && token.text.firstOrNull()?.isUpperCase() == true -> Consequence.MEDIUM // names, places
else -> Consequence.LOW
}With that in place, the interface needs exactly four behaviours: show, underline, ask to confirm, block-and-replay. A low-consequence token never triggers more than an underline, whatever its score. A high-consequence token below the threshold blocks the sentence from being marked as final until the person has replayed it or edited it. Thresholds must be validated on representative recordings; the example policy does not establish a safe operating point.
05
Correction has to cost less than distrust
Figure 03Data flow
Make the correction shorter than the doubt
Read every explanation
- Notice
- Mark a consequential ambiguity where it changes the next action. Avoid displaying a percentage that suggests unverified calibration.
- Listen
- Replay enough context to understand the phrase. Offer alternatives only when the decoder actually supplies them.
- Edit
- Edit the transcript without recapturing the entire phrase. Keep the original audio available to resolve later disagreements.
- Review
- Translate the corrected transcript and review the output. A corrected source word does not prove that the target-language sentence is faithful.
Mark a consequential ambiguity where it changes the next action. Avoid displaying a percentage that suggests unverified calibration.
A cue that leads nowhere trains people to ignore cues. Every marked token in OfflineLingo is tappable: the tap shows the n-best alternatives whisper considered, and a second control replays 1.5 seconds of audio centred on the word. Choosing an alternative or retyping the word updates the translation; the rest of the sentence is not re-run. That last part matters—if fixing one word meant waiting for the whole pipeline again, nobody would fix words.
I also removed the percentage. An early build showed 61 % next to the word and testers spent time reasoning about the number. The underline plus alternatives conveyed the same doubt and led directly to the action. A number invites interpretation; an underline invites a tap.
06
What I would measure next
- Edit rate on flagged versus unflagged high-consequence tokens. If people edit unflagged ones as often, the flags are miscalibrated.
- Time from flag to correction. Above a few seconds, the correction path is too heavy.
- Flag fatigue: how many flags per sentence before people stop tapping. My guess is three; I would rather know.
- How often
compression_ratiocatches a hallucinated segment in real noise. If it is frequent, that check deserves a visible state of its own.
07
A score needs a calibration set
A token probability is conditional on the audio and previously decoded tokens; it is not a calibrated probability that a word is correct. A word may contain several tokens. Validate any aggregation and thresholds on labelled recordings from the intended languages and acoustic conditions. Measure missed consequential errors and unnecessary interruptions separately. A simple token classifier also misses written-out numbers, multiword negations and names without capitals: unknown cases need review, not an automatic low-risk label.
More articles
09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read