Production benchmark · 16 September 2026
Utility bill OCR benchmark
Measured on 51 current production documents matched by exact source hash to reviewer-approved human-gold fixtures. Each bill ran through the production OCR-first recipe without manual correction.
Results
85.9% field accuracy across 824 scored values.
Reviewed corpus
51 documents
Exact source-hash matches to human-gold fixtures
Field accuracy
85.9%
708 of 824 scored field assertions
Processing failures
0 of 51
0.0% runner failure rate
Recipe latency
25.520s median
36.206s p95 using nearest-rank
Field accuracy is correct scored fields divided by all 824 scored fields. Missing expected rows contribute their labeled fields as failures. Extra rows are reported separately.
Latency covers OCR cache lookup or capture, model response, JSON parsing, source binding, validation, reconciliation, and projection.
Method
Production inputs. Independent labels. One frozen run.
Production sources supplied the private population. Existing human-gold fixtures supplied the labels. A document entered the scored subset only when its exact source SHA-256 matched a fixture, so production output never graded itself.
- 01
Inventory the production population
Inventory 2,082 active source-backed placements. Collapse exact source hashes into 2,027 unique documents, and exclude 695 placements without a source hash.
- 02
Bind reviewer-approved ground truth
Match 51 current production documents by exact SHA-256 to existing human-gold fixtures. Production outputs do not grade themselves.
- 03
Freeze the scorer and production revision
Freeze source hashes, labels, normalization rules, scored fields, exclusions, and parser revision d192034a8811 before the run.
- 04
Run the scored subset once
Process each of the 51 documents through the production OCR-first recipe. Keep every field miss, row-count error, and processing failure in the report.
Private corpus
The scored subset came from the production population.
The inventory covered 2,082 active source-backed placements and collapsed exact source hashes into 2,027 unique documents. It excluded 695 placements without a source hash.
Fifty-one current documents matched reviewed fixtures. The other 1,976 remain unscored. Source bills, customer identifiers, raw values, and raw model output remain private.
Field accuracy by utility type
Electricity
21 documents · 315 of 366 scored fields
86.1%
Natural gas
10 documents · 112 of 140 scored fields
80.0%
Water
20 documents · 281 of 318 scored fields
88.4%
Scoring contract
Score fields, not impressions.
Field accuracy equals correct scored fields divided by all scored fields. Missing expected rows contribute their labeled field count as failures. Extra output rows are excluded from that denominator and reported as row-count failures.
Document fields score once per document. Row fields score once per labeled utility row. The table shows exact results from the frozen run.
| Field | Comparison | Correct |
|---|---|---|
| Vendor name | Exact value after frozen scorer normalization | 45 / 51 |
| Customer name | Exact value or approved alternative | 51 / 51 |
| Account number | Exact text | 51 / 51 |
| Invoice date | ISO date | 51 / 51 |
| Due date | ISO date | 43 / 51 |
| Service address | Exact normalized text | 24 / 51 |
| Row: utility type | Declared activity class | 54 / 73 |
| Row: usage | Decimal value | 63 / 73 |
| Row: usage unit | Approved unit mapping | 51 / 73 |
| Row: currency | ISO currency code | 46 / 46 |
| Row: service-period start | ISO date | 72 / 73 |
| Row: service-period end | ISO date | 72 / 73 |
| Row: invoice total | Decimal value | 34 / 46 |
| Row: meter number | Exact text when labeled | 51 / 61 |
Failure policy
Failures stay in the denominator.
The run had no execution failures, but field and structure checks still found errors. Incorrect and missing values remain in the 824-field denominator. Row-count failures stay visible beside field accuracy.
- Missing fields2.2% of 824 scored field assertions
- Row-count failuresExpected and extracted utility rows differed
- Review-gated documentsEvery measured result triggered at least one review gate
- Processing failuresAll 51 documents completed the recipe
Limits
One private production corpus cannot guarantee results for every provider or document.
The scored subset contains 51 of 2,027 inventoried unique documents. The other 1,976 documents do not affect the reported accuracy.
The inventory did not classify digital versus scanned quality. Segment results cover utility type only, not every provider, layout, or geography.
Every measured document triggered at least one review gate. These results do not support unattended use or guarantee performance on a buyer's bill mix.