Three claims that should stay separate
Consider a model that extracts an invoice into valid JSON. The result satisfies a representation rule: it has the expected fields and types. That is useful, but it does not establish that the supplier is the right legal entity, the tax calculation follows the applicable policy, or the invoice has not already been recorded.
Even a result that passes those business checks does not prove a bill exists in the ERP. Representation, business validity, and confirmed effect are separate claims. A reliable interface should expose which claim has been established instead of allowing a single success flag to imply all three.
A result can reconcile and still be wrong
Suppose two supplier records share a trading name. The extracted line items sum correctly and the invoice total matches. A model chooses the first supplier returned by a search. Every arithmetic check passes, yet creating the bill against that supplier would attach the obligation to the wrong entity.
The failure is not solved by a more elaborate explanation of the arithmetic. It requires an identity rule and evidence that distinguishes the records. If the available data cannot do that, the correct outcome is unresolved identity, not a more confident supplier guess.
Put acceptance outside candidate generation
Put a clear boundary between generation and acceptance: a model proposes a candidate; the contract defines the checks that admit that candidate to the authorized action. The model may revise its output after a mismatch, but it should not revise the acceptance rule to make the output pass.
This separation makes model routing meaningful. Different models can attempt the same task under the same requirements. Without a stable acceptance boundary, a fallback might appear to improve success by silently changing the tolerance, omitting a difficult field, or relaxing the definition of completion.
Not every check is deterministic
Some work requires qualitative assessment. A support response may need a rubric for completeness or tone. Treating that judgment as equivalent to a database read or arithmetic equality would overstate its evidence. The result should identify the kind of check and preserve the uncertainty relevant to that claim.
A bounded judgment can still be useful. The contract needs to say what it evaluates, which facts are independently checked, and whether that judgment is sufficient for the permitted action. There is no reason to force every business task into the same verification model.
The implementation consequence is a smaller write surface
An action executor should receive a candidate that has satisfied the required checks, the pinned contract and policy references, and a narrowly authorized operation. It also needs to re-evaluate any critical preconditions that may have changed since verification. Passing an earlier check is not permanent permission to write.
This introduces work: more explicit policy, clearer integration semantics, and more cases that stop for missing evidence. That is a deliberate tradeoff. A system that can explain why it cannot establish completion gives operators a better recovery path than one that converts uncertainty into a successful-looking record.
Missing a detail or found a problem?
Send a documentation question →