Assurance
Evaluation
The framework this system is steered by. Field-level accuracy against a golden set is the floor; operator edit rate is the north star, because it is measured in production, on every filing, at no extra cost — the operator was going to review the form anyway. Their corrections are the label.
Operator edit rate
Field accuracy — exact
character-identical match
Field accuracy — normalised
date and unit forms folded
Auto-accept precision
correct, of those shipped unreviewed
Escaped errors
reached the payer uncorrected
Why two accuracy numbers
Exact match counts 05/04/2026 against
2026-05-04 as a miss. Normalised match folds date, unit and
name-order variants and counts it as a hit. The gap between the two —
— is pure formatting, which is cheap to fix and should never be
confused with a comprehension failure. Reporting only the normalised number hides
real errors; reporting only the exact number invents fake ones.
Operator edit rate by field type
28,221 field decisions · 30 days
This single chart is the roadmap. Direct transfers are solved. Narrative synthesis
is expensive but improving. Quantified hedges — turning "occasional" into a
number — sit at 31.2% and are where the p90 tail lives.
What we would do next, and why
Quantified hedges are only 2.1% of field volume but carry a 31.2% edit rate — they
consume attention out of all proportion to their count, and they are exactly the
fields on the demo case that blew past the budget. The fix is not a better prompt.
It is a structured elicitation path: when a source says "occasional", stop
guessing a number, and instead present the operator a one-keystroke choice between the
two or three defensible readings. That converts a 30-second free-text decision into a
2-second one, and it is worth roughly 1.4 forms per hour — most of the
remaining gap to 15.
Calibration
Reliability by confidence bin
n = 28,422
| Bin | Predicted | Observed | Δ | n | Verdict |
|---|
A confidence score is worthless unless it is calibrated. If the system says 0.90
it must be right about 90% of the time — otherwise the auto-accept threshold is a
guess and every downstream claim rests on it. The
threshold was set from this
table, not chosen because it looked like a reasonable number.
Per-template performance
| Template | Family | Filings | Fields | Edit rate | Accuracy | Auto-accept | p50 | Status |
|---|
Workers' comp templates trail the FMLA family by a wide margin, and it is structural
rather than accidental: state comp forms ask for causation and apportionment
judgements that the clinical record frequently does not contain. That is a
source coverage problem, not a model problem, and no amount of prompt work fixes it —
it needs a different intake conversation with the provider.
Release regression
golden set: 480 forms · 31,204 fields
| Version | Model | Date | Accuracy | Edit rate | Auto-accept | Escaped | Verdict |
|---|
Every prompt or model change is scored against a frozen golden set before it ships.
The 2026-06-02 release was rolled back on the escaped-error
rate alone — accuracy fell less than a point, but errors reaching the payer rose by
more than half, and that is the number with legal consequences.