Choose the unit before calculating a rate
A submission, a run, a model attempt, a verifier result, and a business action are different units. One run may contain several model attempts and many checks. A retried HTTP request may refer to the same run. Counting them interchangeably can make quality and throughput appear better or worse without any change in business behavior.
Start with a cohort of unique accepted runs for a pinned contract version and a defined observation period. Report requests rejected before acceptance separately, including their reasons. Replayed requests under the same operation identity should not create additional successful outcomes in the denominator.
Distinguish candidate quality from completed work
| Measure | Suggested definition | What it does not establish |
|---|---|---|
| Candidate acceptance rate | Candidates satisfying every required check divided by candidates entering verification. | That an authorized action was attempted or confirmed. |
| Confirmed outcome rate | Unique runs fulfilling the contract divided by unique accepted runs in the cohort. | That unfinished runs will never complete. |
| Unavailable-check share | Required check evaluations lacking usable evidence divided by required check evaluations. | That the candidate itself was incorrect. |
| Escalation share | Runs needing another model attempt divided by accepted runs. | That escalation improved correctness or economics. |
Keep unavailable evidence in view
If a receipt service is unavailable, dropping those runs from the reported population can produce a flattering pass rate precisely when the end-to-end service is less useful. A pass rate calculated only among resolved checks can be informative, but it must be labeled conditional and shown alongside the unresolved population.
The same care applies to observation time. Recently accepted runs have had less time to finish. Publish the cohort window, measurement cutoff, and counts still queued, running, reconciling, or awaiting input. A terminal-only success rate is a different metric from completion across all accepted work.
Measure the full path to confirmation
Outcome latency should run from the defined acceptance point to contract fulfillment, including verification, authorized action, and reconciliation where needed. Model response latency is useful for diagnosis but cannot stand in for that end-to-end measure. Report unfinished observations rather than silently assigning them a zero duration.
For economics, account for all relevant attempts, including failed candidates and fallbacks. Dividing total cohort processing cost by confirmed outcomes can expose the cost of unfinished work; when the denominator is zero, the measure is undefined, not free. Processing charges and completion fees should remain distinguishable.
Numbers need a stable contract and a representative workload
A rising completion rate can reflect an easier input population, a narrower contract, relaxed tolerances, or better execution. Keep contract versions, workload characteristics, and policy changes visible before attributing the change to a model or router.
Synthetic fixtures are useful for repeatable failure testing, but they do not establish production performance. Report what was tested, how expected results were determined, and which cases were excluded. This makes comparisons across workloads possible and gives engineers a clear account of the operation’s limits.
Missing a detail or found a problem?
Send a documentation question →