Assurance

Evaluation

The framework this system is steered by. Field-level accuracy against a golden set is the floor; operator edit rate is the north star, because it is measured in production, on every filing, at no extra cost — the operator was going to review the form anyway. Their corrections are the label.

Operator edit rate
Field accuracy — exact character-identical match
Field accuracy — normalised date and unit forms folded
Auto-accept precision correct, of those shipped unreviewed
Escaped errors reached the payer uncorrected
Why two accuracy numbers Exact match counts 05/04/2026 against 2026-05-04 as a miss. Normalised match folds date, unit and name-order variants and counts it as a hit. The gap between the two — — is pure formatting, which is cheap to fix and should never be confused with a comprehension failure. Reporting only the normalised number hides real errors; reporting only the exact number invents fake ones.
Operator edit rate by field type
28,221 field decisions · 30 days
This single chart is the roadmap. Direct transfers are solved. Narrative synthesis is expensive but improving. Quantified hedges — turning "occasional" into a number — sit at 31.2% and are where the p90 tail lives.
What we would do next, and why Quantified hedges are only 2.1% of field volume but carry a 31.2% edit rate — they consume attention out of all proportion to their count, and they are exactly the fields on the demo case that blew past the budget. The fix is not a better prompt. It is a structured elicitation path: when a source says "occasional", stop guessing a number, and instead present the operator a one-keystroke choice between the two or three defensible readings. That converts a 30-second free-text decision into a 2-second one, and it is worth roughly 1.4 forms per hour — most of the remaining gap to 15.
Calibration
Reliability by confidence bin
n = 28,422
BinPredictedObserved ΔnVerdict
A confidence score is worthless unless it is calibrated. If the system says 0.90 it must be right about 90% of the time — otherwise the auto-accept threshold is a guess and every downstream claim rests on it. The threshold was set from this table, not chosen because it looked like a reasonable number.
Per-template performance
TemplateFamilyFilingsFields Edit rateAccuracy Auto-acceptp50Status
Workers' comp templates trail the FMLA family by a wide margin, and it is structural rather than accidental: state comp forms ask for causation and apportionment judgements that the clinical record frequently does not contain. That is a source coverage problem, not a model problem, and no amount of prompt work fixes it — it needs a different intake conversation with the provider.
Release regression
golden set: 480 forms · 31,204 fields
VersionModelDate AccuracyEdit rate Auto-acceptEscapedVerdict
Every prompt or model change is scored against a frozen golden set before it ships. The 2026-06-02 release was rolled back on the escaped-error rate alone — accuracy fell less than a point, but errors reaching the payer rose by more than half, and that is the number with legal consequences.