All datasets Document AI

OCR Ground-Truth Document Factory

We generate reproducible, visibly synthetic invoice-shaped test pages together with exact OCR and structured ground truth: word and line boxes, reading order, typed entity fields, table cells, and reconciled amounts, all produced by the renderer itself rather than a human annotator or a second OCR engine. One output record is one generated page: its metadata, token and entity counts, SHA-256 hashes, and the keys to its PNG, raster PDF, and full annotation JSON in the run's key-value store.

Refresh
Generated on demand
Price
$50 per 1,000 pages

Who uses this

OCR and IDP model evaluation teams

Measure word error rate, box IoU, reading order, and field extraction accuracy against exact, renderer-produced labels.

Pipeline and regression testing teams

Feed a document pipeline a fixed, reproducible corpus in CI instead of a folder of real customer files.

Benchmark and fixture set builders

Build a labeled evaluation set of any size without collecting, redacting, or licensing real documents.

Robustness testing teams

Compare clean renders against deterministic scan degradation at the same underlying geometry.

Sample pages

Example pages from a run of the generator, shown with a subset of the fields. Every run returns the full field set below.

documentIdseedtemplateIdprofiledpitokenCount
SYN-000DBBA1-0000900001invoice_compact_v0clean144118
SYN-000DBBA2-0001900002invoice_grid_v0scan_light200123
SYN-000DBBA3-0002900003invoice_sidebar_v0scan_hard300100

Example records from three pages generated with the same inputs as the actor's recorded smoke run (seeds 900001 to 900003).

Fields

The most used fields in each page. The Apify listing documents the full output schema.

documentId
Unique id of the generated page, e.g. SYN-0000002D-0003.
seed
The deterministic seed that produced this exact page; re-running it reproduces byte-identical output.
templateId
Which of the four structural layouts was used: classic, compact, grid, or sidebar.
profile
Scan profile applied: clean, scan_light, or scan_hard.
dpi
Page resolution: 144, 200, or 300.
tokenCount
Number of OCR word tokens on the page, from the annotation ground truth.
entityCount
Number of typed entity fields on the page (invoice number, dates, vendor/customer blocks, subtotal, tax, total, line items).
recordKeys
Key-value store keys for this page's PNG, raster PDF, and annotation JSON artifacts.

How it works

1

Run it on Apify

Open the listing, sign in to Apify, and press Start. The actor reads the public source directly, normalizes each page, and writes the results to your Apify dataset. You can also schedule it to run on a cadence.

2

Filter to what you need

  • seed: required starting seed, 0 to 2,147,483,647; page n of the run uses seed + n.
  • count: pages to generate this run, 1 to 16, default 1; each page is billed once.
  • template_id: auto or one of four structural layouts (Classic, Compact, Grid, Sidebar), default auto.
  • profile: clean, scan_light, or scan_hard deterministic scan degradation, default clean.
  • dpi: exactly 144, 200, or 300, default 144; any other value is rejected.
  • schema_version: pinned output contract, default and only accepted value invoice-ground-truth/v0.

3

Export or connect

Download the dataset as CSV, JSON, or Excel from the Apify console, or pull it straight into your own system through the Apify API and integrations.

1 to 16 pages per run; generate a larger corpus by launching several runs with non-overlapping seed ranges. Single-page English USD invoice-shaped fixtures only, across four structural layouts, three scan profiles, and three DPI settings. Labels are renderer-owned, not human-reviewed: they describe exactly what was drawn, not a real-world invoice distribution.

Pricing

$50 per 1,000 pages

Pay per result, billed through your Apify account at $0.05 per page. You are charged only for the pages the run delivers to your dataset, so a filtered run that matches nothing costs nothing. No subscription, no minimum.

Source, refresh, and licence

Refresh
Generated on demand
Licence
Not applicable in the usual third-party sense: every page is wholly generator-produced synthetic content, not sourced or scraped data. All names, addresses, identifiers, and amounts are invented; no caller-supplied document text, logos, addresses, images, URLs, or templates are accepted. Every page is visibly marked SYNTHETIC TEST DOCUMENT - NOT PAYABLE, and pages must not be used for billing, payment requests, identity construction, deception, or fraud.

Questions buyers ask

What exactly do I get for each generated page?

A page image (PNG) at 144, 200, or 300 DPI, a raster-only PDF containing the same final pixels, and an annotation JSON with tokens, lines, typed entities (invoice number, dates, vendor and customer blocks, subtotal, tax, total), table cells, the reconciled invoice record, and page geometry, all keyed off the dataset item's recordKeys.

Is this real invoice or company data?

No. All names, addresses, identifiers, and amounts are invented, every page is visibly marked SYNTHETIC TEST DOCUMENT - NOT PAYABLE, and the Actor accepts no caller-supplied document text, logos, addresses, images, URLs, or templates.

Will the same input always produce the same output?

Yes. A given seed, template_id, profile, dpi, and Actor version always produce byte-identical output, verifiable against the SHA-256 hashes in each dataset item, so an evaluation set can be described entirely by its input rather than stored and shipped around.

Pull the data you need

Run it yourself on Apify, or ask us for a custom extract, a join against another dataset, or a scheduled delivery.