A model choice leaves the important questions open
“Use the strongest model for invoice processing” says little about the operation. Does processing mean extraction, reconciliation, draft creation, accounting posting, or payment? Which source establishes supplier identity? What happens when the purchase order is missing?
Those questions determine correctness and permission. They remain unanswered even if a model performs well on a document benchmark. A model can extract every field accurately while the surrounding workflow creates the bill in the wrong company or writes it twice after a timeout.
Write the promise at the business boundary
A useful contract names the accepted input, required context, output schema, mandatory checks, permitted action, and evidence for completion. It also names exclusions. Creating a draft should not inherit permission to approve or pay merely because those actions are available in the same ERP.
A narrower first contract trades apparent coverage for a clearer result. Supporting one supplier identity model and a defined purchase-order workflow can produce a testable operation. Claiming to handle every invoice format and exception before those branches are specified makes the word complete difficult to interpret.
Route inside the promise
Once the finish line is fixed, routing becomes an implementation choice constrained by data policy, spending limits, latency needs, and evidence from prior attempts. A lower-cost candidate can be useful if it passes the same required checks. A stronger fallback can repair a failed extraction without changing the business policy.
The tradeoff is that some tasks are difficult to verify cheaply. If a qualitative judgment dominates acceptance, route evaluation needs an appropriate assessment method and a clear account of its limitations. The contract does not magically make every model comparison objective.
A contract change can be a permission change
Suppose a contract version originally creates a draft catalog record. A later version publishes the record immediately. The output fields may barely change, but the business effect and risk do. Treating this as an invisible implementation update would violate a caller’s reasonable understanding of the operation.
Versions should therefore capture meaning as well as shape. Verifier requirements, applicability rules, action semantics, and evidence expectations need traceable changes. A historical receipt should remain interpretable against the rules that actually governed its run.
Evaluate the operation before expanding the library
A useful test set includes invalid inputs, missing context, ambiguous identities, changed source records, concurrent duplicate submissions, exhausted budgets, and interrupted writes. The question is not only whether a candidate looks correct; it is whether the operation produces the right effect and explains cases it cannot complete.
This approach favors a small number of well-defined capabilities over a long list of loosely described tasks. The number of contract names is not evidence of useful coverage. Coverage becomes meaningful when each boundary, limitation, and recovery path can be inspected and exercised.
Missing a detail or found a problem?
Send a documentation question →