One timeout, two possible histories
A draft-bill request can time out before the ERP receives it. It can also time out after the ERP creates the bill but before the response reaches the caller. Those histories look similar to the client and require different recovery actions.
A blind retry is helpful in the first history and may create a duplicate in the second. Calling the run failed does not resolve the ambiguity. Asking another model to generate the bill again is even less relevant: the uncertainty concerns a system effect, not the quality of the extracted invoice.
Preserve the identity of the work
Recovery needs a stable operation identifier that survives transport attempts. Binding an idempotency key to the original request and run protects the submission boundary. The destination write still needs its own reliable deduplication or lookup mechanism.
An HTTP method being idempotent describes the intended effect of repeated identical requests. It does not establish that an arbitrary business operation executes exactly once across every downstream system. The connector must account for the actual destination semantics, including uniqueness constraints, native idempotency, and asynchronous jobs.
Reconciliation is useful work
After a lost response, the runtime should look for the original operation through an authoritative reference. If it finds the expected draft and confirms the relevant values, it can establish completion without writing again. If it establishes that no operation occurred, a controlled retry may be possible.
Absence from a stale read replica is weaker evidence than absence from the authoritative write system. The connector needs to define which query can settle the question and how visibility delays are handled. A bounded reconciliation window may end with uncertainty still unresolved.
{
"sample": true,
"state": "reconciling",
"verification_status": "passed",
"effect_status": "unknown",
"operation_id": "sample_operation_001"
}Some destinations cannot support the same promise
A destination without a unique external reference, a reliable search path, or native idempotency may not permit safe automatic recovery. Local locking can reduce concurrent attempts inside one runtime, but it cannot discover an unobservable remote commit after a failure.
That limitation should affect the contract and integration policy. The system may need to stop and request a reconciliation decision. Advertising universal exactly-once business execution would conceal a property that depends on each destination and operation.
Cancellation and billing follow the same distinction
Cancellation can stop future work, but it cannot erase an operation already accepted by another system. Reversing a confirmed action is a separate business operation with its own permissions and consequences. It should not be disguised as a successful cancellation.
Similarly, confirmed completion and incurred processing are different accounting facts. A fee tied to completion needs evidence that the contract was fulfilled; the runtime reaching a timeout cannot supply that evidence. The receipt should retain the uncertainty, the reconciliation observations, and the eventual conclusion when one becomes available.
References
Missing a detail or found a problem?
Send a documentation question →