Parsepoint
Menu

Production benchmark · 16 September 2026

Utility bill OCR benchmark

Measured on 51 current production documents matched by exact source hash to reviewer-approved human-gold fixtures. Each bill ran through the production OCR-first recipe without manual correction.

Results

85.9% field accuracy across 824 scored values.

Reviewed corpus

51 documents

Exact source-hash matches to human-gold fixtures

Field accuracy

85.9%

708 of 824 scored field assertions

Processing failures

0 of 51

0.0% runner failure rate

Recipe latency

25.520s median

36.206s p95 using nearest-rank

Field accuracy is correct scored fields divided by all 824 scored fields. Missing expected rows contribute their labeled fields as failures. Extra rows are reported separately.

Latency covers OCR cache lookup or capture, model response, JSON parsing, source binding, validation, reconciliation, and projection.

Method

Production inputs. Independent labels. One frozen run.

Production sources supplied the private population. Existing human-gold fixtures supplied the labels. A document entered the scored subset only when its exact source SHA-256 matched a fixture, so production output never graded itself.

  1. 01

    Inventory the production population

    Inventory 2,082 active source-backed placements. Collapse exact source hashes into 2,027 unique documents, and exclude 695 placements without a source hash.

  2. 02

    Bind reviewer-approved ground truth

    Match 51 current production documents by exact SHA-256 to existing human-gold fixtures. Production outputs do not grade themselves.

  3. 03

    Freeze the scorer and production revision

    Freeze source hashes, labels, normalization rules, scored fields, exclusions, and parser revision d192034a8811 before the run.

  4. 04

    Run the scored subset once

    Process each of the 51 documents through the production OCR-first recipe. Keep every field miss, row-count error, and processing failure in the report.

Private corpus

The scored subset came from the production population.

The inventory covered 2,082 active source-backed placements and collapsed exact source hashes into 2,027 unique documents. It excluded 695 placements without a source hash.

Fifty-one current documents matched reviewed fixtures. The other 1,976 remain unscored. Source bills, customer identifiers, raw values, and raw model output remain private.

Field accuracy by utility type

  • Electricity

    21 documents · 315 of 366 scored fields

    86.1%

  • Natural gas

    10 documents · 112 of 140 scored fields

    80.0%

  • Water

    20 documents · 281 of 318 scored fields

    88.4%

Scoring contract

Score fields, not impressions.

Field accuracy equals correct scored fields divided by all scored fields. Missing expected rows contribute their labeled field count as failures. Extra output rows are excluded from that denominator and reported as row-count failures.

Document fields score once per document. Row fields score once per labeled utility row. The table shows exact results from the frozen run.

FieldComparisonCorrect
Vendor nameExact value after frozen scorer normalization45 / 51
Customer nameExact value or approved alternative51 / 51
Account numberExact text51 / 51
Invoice dateISO date51 / 51
Due dateISO date43 / 51
Service addressExact normalized text24 / 51
Row: utility typeDeclared activity class54 / 73
Row: usageDecimal value63 / 73
Row: usage unitApproved unit mapping51 / 73
Row: currencyISO currency code46 / 46
Row: service-period startISO date72 / 73
Row: service-period endISO date72 / 73
Row: invoice totalDecimal value34 / 46
Row: meter numberExact text when labeled51 / 61

Failure policy

Failures stay in the denominator.

The run had no execution failures, but field and structure checks still found errors. Incorrect and missing values remain in the 824-field denominator. Row-count failures stay visible beside field accuracy.

  • Missing fields2.2% of 824 scored field assertions
  • Row-count failuresExpected and extracted utility rows differed
  • Review-gated documentsEvery measured result triggered at least one review gate
  • Processing failuresAll 51 documents completed the recipe

Limits

One private production corpus cannot guarantee results for every provider or document.

The scored subset contains 51 of 2,027 inventoried unique documents. The other 1,976 documents do not affect the reported accuracy.

The inventory did not classify digital versus scanned quality. Segment results cover utility type only, not every provider, layout, or geography.

Every measured document triggered at least one review gate. These results do not support unattended use or guarantee performance on a buyer's bill mix.