Consider a workflow that creates a support ticket. The service saves the ticket, but the connection drops before the caller receives its identifier. This design note follows that illustrative failure through recovery, concurrency and testing.
Explore this article visually2 figures
01
The missing reply is an unknown outcome
The tempting recovery is to run the same create call again. If the first request already committed, that produces a second ticket. If it never arrived, a retry may be exactly what is needed. The caller sees the same timeout in both cases.
Figure 01Comparison
Two histories, one timeout
Read every explanation
- Never committed
- The request never reached a successful commit. A new attempt may complete the original intent.
- Committed, reply lost
- The service created the ticket. Repeating an unprotected create operation can add another one.
The request never reached a successful commit. A new attempt may complete the original intent.
Keep “unknown” distinct from “failed” in the interface and stored state. A useful message is “Checking whether the ticket was created”, accompanied by an operation reference. A success message without a confirmed result would conceal the very uncertainty the recovery process must resolve.
02
Name the intent before making the request
For this proposed design, create an operation record before dispatch. It contains a caller scope, operation key, request fingerprint and state. A second attempt for the same intent reuses that record. A person deliberately creating another ticket gets a new key, even when the text is identical.
Figure 02Data flow
One operation survives multiple attempts
Read every explanation
- Persist intent
- Save the identity before sending. Recovery after a process restart must find the same key.
- Dispatch
- Pass the key through the service’s supported idempotency mechanism. A local identifier alone cannot protect a remote write.
- Resolve
- Store the returned ticket ID, or keep the operation unresolved until the service can establish its outcome.
Save the identity before sending. Recovery after a process restart must find the same key.
Idempotency makes repeated attempts represent the same operation within a defined contract. AWS describes client request identifiers for this purpose; Stripe documents a concrete API implementation. Check the service’s scope, retention and parameter-matching rules before relying on its guarantee.
03
Design the collision and crash paths
In a service you own, enforce uniqueness on the caller scope and operation key. Compare the stored request fingerprint before replaying a result. The same key with different parameters is a conflict to surface, not permission to overwrite the original intent.
Two workers can receive the same operation at once. Reserve ownership atomically and define how an in-progress response behaves. Saving “complete” after the write leaves a crash window; saving it before the write can claim a result that does not exist. When both records live in one database, a transaction can tie the business write to its operation result.
A local database transaction cannot atomically commit an unrelated external API call. If the remote system offers neither idempotency nor a reliable lookup by operation reference, automatic recovery cannot promise duplicate prevention. Keep the operation unresolved and route it to reconciliation. Reconciliation must establish what happened, not merely wait and assume failure.
04
Spend a retry budget deliberately
Retry only failures the service contract considers recoverable. Use bounded backoff with jitter so clients do not all retry together. AWS’s guidance explains why retries can increase load on an already struggling dependency.
Choose one layer to own the retry budget. If three layers each allow three total attempts, a single top-level request can cause up to 27 downstream attempts. Count SDK retries as well as application retries. Set an overall deadline that includes waiting and execution; do not start another attempt after the remaining time is insufficient.
| Observed state | Next action | Keep visible |
|---|---|---|
| Confirmed success | Return the saved result | Original resource ID |
| Recoverable rejection | Retry within the contract and budget | Attempt count and next check |
| Invalid request | Correct the input | Specific validation error |
| Timeout after dispatch | Reconcile or retry with duplicate protection | Unknown outcome |
| Protection window expired | Establish the original outcome first | Manual review if unresolved |
05
Test the response you never received
A useful fault-injection test commits the ticket and then drops the response. Restart the worker and replay the operation with its original key. The assertion is one business resource and a consistent recovered identifier, not merely a successful HTTP response.
- Send the same key concurrently from two workers; verify one intended effect.
- Reuse a key with a changed payload; expect a conflict.
- Crash before dispatch and after remote commit; check both recovery paths.
- Exhaust the deadline; verify that no new attempt starts afterward.
- Simulate expired duplicate protection; verify reconciliation rather than a blind create.
Track unresolved operation age, duplicate effects and attempts per completed intent. A high eventual-success rate can hide a growing queue of uncertain writes. The recovery path is complete only when the system can explain which action happened and connect it to the original request.
More articles
08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read02Designing AI that works without a networkEdge AI · 8 min read