EnviroBench

Methodology: tasks, metrics, evaluation protocol and running the benchmark

Version
0.2.0 (pre-release)
Benchmark types
17
Runtime
Ruby 3.3, standard library only
Repository
Private until release

EnviroBench is not yet public. The code runs today on small synthetic datasets; the public development set, sealed holdout and license are still in preparation (§7). This page describes version 0.2.0 as implemented and marks planned work as planned.

Abstract

EnviroBench tests AI systems on document tasks that environmental teams do: reading analytical table cells, mapping table cells to samples and measurements, locating tables and their row and column lines on scanned pages, placing sample locations on site plans, deciding which OCR readings can skip human review, and ranking passages for search. Each benchmark type fixes an input format, an answer format and deterministic per-case scoring. The system under test runs behind an adapter that receives inputs only. EnviroBench joins the held-back answers afterwards, records every error on every case, and compares a run case by case with a saved baseline and same-code control runs. A change passes only when it introduces no new error on a case that is stable across repeats; a better average never cancels a new error.

1Design

EnviroBench separates what is measured from what is being measured. The public core defines benchmark types, answer formats, scoring rules, the adapter protocol and the report. It does not run any product or model. Datasets come from plugins, and each system under test is an adapter. An organization can keep its own plugin, adapter and object store private while using the same public scoring.

Table 1. Components and what each one owns.
ComponentOwnsNever does
Core Benchmark types (one Ruby file each), the run lifecycle, scoring, pairing, gates, review drafts and the HTML report. Public synthetic datasets. Run a product, call a model, or hold production-derived data.
Plugin Datasets: cases, answer keys, checksummed source files, optional source links, and saved run profiles (bench.json). Define scoring. Plugins add data only.
Adapter Runs one system or model on input-only cases and returns its predictions, cost and usage. See answers, earlier runs or source links, or score itself.

2Tasks and metrics

A benchmark type is one class in lib/envirobench/benchmarks/. It declares an identifier, a result format, a scoring version and its score profiles, and implements score_case, which returns, for each profile, the case's errors and metrics. Errors are comparable values, so the same mistake in two runs is recognized as the same error.

Table 2. Benchmark types in version 0.2.0. Coordinates use 0–1000 units across the image unless stated. The synthetic column counts cases in the public smoke dataset.
Benchmark Input Prediction Unit Synthetic
document-extraction One whole PDF (digital or scanned) and a statement of which values are in scope Every field-sample measurement as a record: location, sample, date, depth, matrix, parameter, result, unit, non-detect and limit, qualifiers document 40
site-dataset A site's whole document record: every PDF, most pages holding no results The site's deduplicated analytical dataset, each record citing a page that shows its governing value site 1
site-questions The same record and a fixed question catalogue Typed, cited answers, or "not in the record" site 1
sample-location-placement Site plan image and a sample-location label Whether the marker was found, and its point in percent of the image point 24
table-cell-ocr Crop of one table cell and the existing first-pass reading The cell's text cell 80
table-mapping A table as CSV, with optional cell boxes Orientation, a role for each cell (sample identity, chemical name, value, unit, method), value groups and default units table 37
table-separators Table crop image, with optional word boxes Positions of vertical and horizontal separator lines crop 13
table-containers Page image, with optional detector boxes and word boxes One box per table, and an optional angle that straightens the page page 11
ocr-auto-accept Saved readings of one cell from several OCR and vision readers, plus image checks Accept one value, or send the cell to review cell 100
search-reranking A query and 1–100 retrieved passages An ordering of passage IDs query 46
instrument-execution Page images of a legal instrument and any related parent, schedule or reference pages, with its title, type and the parties expected to sign Execution status (executed, partial, unsigned, unknown, missing, referenced only), the parties visibly signed and the execution page instrument 51
party-roles Page images of one document (tank registration, lease, consignment agreement, title, permit, notice or letter), its title and type, and the parties it names The roles the document states for each party: owner, tank owner, operator, registrant, lessor, lessee, consignor, consignee or other document 42
indemnity-terms A contract's definitions, indemnity and release clauses, limitations, survival, schedules and amendments as page images, with its title, type and parties Each indemnity and release: kind, giver and beneficiary, environmental coverage, period before or after closing, cap, survival and carve-outs contract 41
chain-of-title A site's title documents (certificates, searches, abstracts) as page images, including a neighbouring parcel's search, and the subject parcel's legal description The subject parcel's ownership periods: owners, start and end dates as printed, granting instrument, consideration, and whether a period continues the same owner under a new name chain 40
rosc-fields An AER Record of Site Condition form (OneStop Submission PDF); a public real-world set of 84 forms downloads from AER at pinned URLs Submission ID and date, licensee and its identifier, site name, intents, and the UWIs and related submissions listed, each with a quote form 4
source-conflicts A site record of 7–10 documents (assessments, laboratory, regulator, closure, tank, title and spill records) as page images, with a document list Each material conflict: fact type, both sides cited by document and page, and the governing source where a rule settles it site record 30
missing-information A site record with documents withheld, and questions from a fixed catalogue (owner, tanks, last groundwater event, remediation, regulatory status, spills, latest exceedance) Each answer with a citation, or not in the record; the documents the record cites but does not hold site record 30

2.1Errors and headline metrics

Each profile is a separate way of scoring the same prediction. The first profile is the default gate. Headline metrics summarize a run; the gate uses only per-case errors.

Table 3. Score profiles, what counts as an error, and headline metrics.
Benchmark Profile A case has an error when Headline metrics
document-extraction records An expected record has no predicted match, or a predicted record matches nothing. Records match on parameter (CAS number, a public synonym table, or the name) and sample identity, then by an optimal one-to-one assignment. Correct-record F1, pooled over all documents: 2 × matched records whose value is right (non-detect state, number or limit, unit, and reporting limit where the answer gives one) ÷ (expected + predicted records). Record F1, counting every matched record whatever its value.
values On a matched record, the result (after unit conversion), non-detect state, unit, reporting limit or qualifiers differ. A wrong value with no non-detect, qualifier or review flag is also a silent wrong value. Value accuracy (non-detect state, number and unit right); silent wrong values as a share of matched records, lower is better.
metadata On a matched record, the location, sample ID, date, time, depth or matrix differ. Field accuracy per field.
site-dataset records A measurement is missing, invented, or reported again (a duplicate of one already reported). Correct-record F1 (right governing value and a valid citation); duplicate rate, lower is better; record F1.
values As document extraction, plus a superseded value: a number a correction letter or a final report replaced. Silent wrong values, lower is better.
metadata As document extraction, plus a citation to a page that does not show the governing value. Valid citations.
site-questions default An answer is wrong or missed, an answer is given where the record has none (invented, severe), or a right answer cites no supporting page (severe). Questions right; invented answers and unsupported citations, lower is better; data questions right.
sample-location-placement default The point is more than 1% of the image diagonal from the checked location: near (within 5%), far, outside the image, or the marker was not found. Points within 1% of the diagonal; within 5%; labels found; median error as % of diagonal.
table-cell-ocr default The reading differs from the checked text after whitespace normalization; a wrong reading that equals the first-pass value is also a false accept, because it would skip review. Cells read exactly; wrong auto-accepts as a share of auto-accepted readings; share auto-accepted.
table-mapping exact Any orientation, cell role, value group or default unit is missing or extra. Tables matching the expected mapping exactly.
important An important fact differs after a plugin-supplied projector turns the mapping into records (for example, the measurements a product would import). Available important checks; measurement correctness; sample assignment.
table-separators strict An expected line has no predicted line within 5 units (matched nearest-first, one to one), or a predicted line matches nothing. Line F1 pooled over all crops and both axes; crops with every line right; line edits per crop.
logical As strict, after replacing each line with the number of word centres before it, so lines that split the same words are equal. Crops splitting the words correctly.
word pairs Words are grouped into different rows or columns than in the expected grid. Word-pair F1, rows and columns averaged.
table-containers boxes An expected box has no predicted box at IoU ≥ 0.3, a predicted box matches nothing, or a matched box has an edge more than 10 units off. Pages needing no box edit; mean IoU; missed and extra boxes per page.
angle The straightening angle is more than 0.3° from the expected angle. Pages straightened within 0.3°.
ocr-auto-accept default An accepted value differs from the checked value, the cell was not fully legible, or blank and non-blank are confused. Sending a cell to review is never an error. Legible cells auto-accepted correctly; wrong auto-accepts; cells sent to review per page.
search-reranking default A judged relevant passage is outside the top 10, or a judged irrelevant passage is inside it. For the gate, a query regresses when nDCG@10 falls by more than 0.01 or fewer essential passages stay in the top 10. nDCG@10 (graded gain, log2 discount); relevant passages in the top 10; queries keeping every essential passage.
instrument-execution default The status differs (an instrument called executed when not fully executed, and an unreadable, absent or only-referenced execution page called unsigned, are their own errors), a signed party is missed or an unsigned one listed, or the execution page differs. Statuses correct; instruments called executed when not; evidence gaps called unsigned; signatory recall and precision.
party-roles default A party is given a role the document does not state for it (an owner or tank-owner role is its own error) or is missing a stated role. Parties with every role right; non-owners called owners; stated roles found; predicted roles the document states.
indemnity-terms default A provision is missing or invented, or its environmental coverage, period, cap, survival or carve-outs differ; environmental coverage claimed where the contract does not give it, and a dropped environmental carve-out, are their own errors. Contracts read exactly right; contracts overstating environmental protection; provisions found; provision fields right; carve-outs found.
chain-of-title default A period is missing, or its instrument, dates, consideration or continuation differ; a period from the neighbouring parcel, and an owner the parcel never had, are their own errors. Periods exactly right; periods from another parcel or invented; owners found; dates right as printed; chains exactly right.
rosc-fields default A field's value differs from the form (identifiers compared without spacing), a value is missed, a list item is missed or added, or a blank field is filled in. Fields read right; forms read exactly right; blank fields filled in; values with a matching quote.
source-conflicts default A conflict is missed or cited on the wrong page, or a non-conflict is reported (restatements, replaced drafts, other sampling events and a neighbour's facts are their own kinds); fact type and governing source are also checked. Conflicts found and cited; conflicts found; false conflicts; both sides cited right.
missing-information default An answer is wrong, missed or unsupported by its citations, a question the record cannot answer is answered (an invented answer), or an absent document is missed or misreported. Questions right; invented answers; answerable questions right; absent documents found.

Scoring versions are recorded in every run, for example envirobench.table-separators.scores.v1. A change to a formula gets a new version.

2.2Tiers

Each benchmark type declares a tier: what a case hands the system. The component tier hands it a crop, a cell or a page; the document tier one whole PDF; the site tier a site's whole document record, a few hundred pages of reports, certificates, letters, invoices and scans, most of them holding nothing that is asked for. Real work starts at the site tier: finding the pages that matter, reading them, and reconciling the same results reprinted across years of reports, drafts and finals, scanned copies and correction letters. The lower tiers show where a system loses accuracy on the way.

At the site tier cost and wall time per site are part of the result, reported beside every score: a system that is accurate only at a high price per site is a different product from one that is accurate cheaply, so site boards sort by both. Three public sites (a service station, a bulk plant on railway lands and a dry cleaner, from 275 to 806 pages) are invented and generated with their answers; real site records stay in private plugins.

2.3Category scores

Leaderboards summarise the tasks in six categories: Extract, Verify, Locate, Research, Legal and Site. Each component task is first turned into a score from 0 to 100, higher better. By default the component score is 100 × the task's primary headline measure; every current primary measure is a share or a mean between 0 and 1 where higher is better, and a lower-is-better primary would need its own published formula. Two tasks built around one dangerous error are the exceptions. OCR auto-accept scores 100 × max(0, (correct accepts − 10 × wrong accepts) ÷ readable cells), so one wrong value that would skip review cancels ten correct accepts and coverage cannot buy back a wrong accept cheaply. Instrument execution scores 100 × max(0, (correct statuses − 5 × instruments called executed when not) ÷ instruments), so a release or transfer wrongly reported as signed cancels five correct statuses, and calling every instrument executed scores 0. Party roles scores 100 × max(0, (parties with every role right − 5 × non-owners called owners) ÷ parties), the same weight for a party wrongly made an owner, and indemnity terms 100 × max(0, (contracts read exactly right − 5 × contracts overstating environmental protection) ÷ contracts), and chain of title 100 × max(0, (periods exactly right − 5 × periods from another parcel or invented) ÷ periods). In Research, missing information scores 100 × max(0, (questions right − 5 × invented answers) ÷ questions).

A category score is the weighted mean of its component scores, Σ wisi ÷ Σ wi. Weights are per task, not per case, so a large test set does not outweigh a small one. Every weight is 1 except document extraction in Extract.

Document extraction is Extract's end-to-end task: a whole PDF in, measurement records out. It is the outcome environmental professionals actually use, so it has weight 4, as much as the four table tasks together (4 of the category's 9). The table tasks show where a system loses records on the way. Its primary measure is correct-record F1: a record counts only when it is found and its value is right.

  • Complete or nothing. A system gets a category score only when it has results for every component task, on the test-set versions the leaderboard shares. Otherwise the category shows Incomplete (n of m tasks). A missing task is never counted as zero and never dropped.
  • Repeats. With repeated runs, the score is the weighted mean of each task's mean, shown with a range: the weighted mean of each task's lowest and highest repeat.
  • No overall score. The categories cover very different amounts of work, so they are not combined into one number.

The definition lives in core (lib/envirobench/categories.rb) with its own scoring version; the table below is generated from it.

Table 4. Categories, their component tasks, weights and component scores. Scoring version envirobench.categories.v6.
CategoryComponent taskTierWeightComponent score (0–100)
EnviroBench ExtractHow accurately does the system turn documents into data?document-extractiondocument4100 × correct-record F1
table-containerscomponent1100 × pages needing no box edit
table-separatorscomponent1100 × line F1 (all lines)
table-cell-ocrcomponent1100 × cells read exactly
table-mappingcomponent1100 × exact mapping agreement
rosc-fieldsdocument1100 × fields read right
EnviroBench VerifyDoes the system know when it might be wrong?ocr-auto-acceptcomponent1100 × max(0, (correct accepts − 10 × wrong accepts) ÷ readable cells)
EnviroBench LocateCan the system place things on site plans?sample-location-placementcomponent1100 × within 1% of the image diagonal
EnviroBench ResearchDoes the system find the right evidence, and see where the record conflicts or falls short?search-rerankingcomponent1100 × nDCG@10
source-conflictssite1100 × conflicts found and cited
missing-informationsite1100 × max(0, (questions right − 5 × invented answers) ÷ questions)
EnviroBench LegalDoes the system read legal instruments without overstating them?instrument-executiondocument1100 × max(0, (correct statuses − 5 × called executed when not) ÷ instruments)
party-rolesdocument1100 × max(0, (parties with every role right − 5 × non-owners called owners) ÷ parties)
indemnity-termsdocument1100 × max(0, (contracts read exactly right − 5 × contracts overstating environmental protection) ÷ contracts)
chain-of-titlesite1100 × max(0, (periods exactly right − 5 × periods from another parcel or invented) ÷ periods)
EnviroBench SiteCan the system work from a site's whole document record?site-datasetsite1100 × correct-record F1
site-questionssite1100 × questions right

3Evaluation protocol

One run covers one benchmark and one system. Figure 1 shows the lifecycle. The boundary between core and adapter is the only place data crosses, and answers never cross it.

Figure 1. Run lifecycle Core loads the dataset and writes an input-only request. The adapter runs the system and returns candidate predictions. Core validates them, joins the held-back answers and saved runs, scores every case per profile, then compares with the baseline and same-code controls to decide the gate and write the report. EnviroBench core Adapter (system under test) 1Load dataset cases, answers, checksummed files 2Write input-only request no answers, earlier runs or links 3Run the system candidate predictions, cost 4Validate and join answers + baseline + controls 5Score every case errors and metrics per profile 6Pair case by case new errors; unstable cases masked 7Gate pass, fail or baseline required report.html · summary.md · metrics.json · review-brief.json predictions only
  1. CoreLoad dataset. Cases, answers and checksummed source files.
  2. CoreWrite an input-only request. No answers, earlier runs or source links.
  3. AdapterRun the system. Return candidate predictions and cost.
  4. CoreValidate and join. Add answers, the baseline and controls.
  5. CoreScore every case. Errors and metrics per profile.
  6. CorePair case by case. New errors; unstable cases masked.
  7. CoreGate and report. Pass, fail or baseline required.
Figure 1. Run lifecycle. Core never modifies the adapter's files; a malformed, missing or duplicate prediction stops the run before anything is scored.

3.1Answer isolation

The adapter is started as ADAPTER run --request FILE --output DIR. The request holds the run ID, model, spend limit, adapter options and an input-only view of the dataset: cases and a run-specific directory of checked source files. It never contains answer rows, the answer file's path, earlier runs' predictions or source links. The adapter returns only the candidate's predictions. Core checks each one with the type's validator; a technical failure blocks the report (exit status 1) and is never counted as a score.

3.2Paired comparison and the regression gate

A run is compared with saved runs rather than with a second model in the same run:

  • --baseline-run DIR is the run the change is compared with.
  • --control-run DIR, repeatable, is a same-code repeat of the baseline. A case is unstable in a profile when the baseline and any control disagree. Its regressions are reported but do not fail the gate. More controls catch more noise.
  • --target-case ID names cases the change must get right. Any remaining error on a target fails the gate.

Per profile, a case regresses when the candidate has an error the baseline did not have. The gate passes when no stable case regressed and every target is error-free. Without a baseline the status is baseline required; when a profile cannot score every case and run, it is unavailable. A type may redefine when a case is worse: search reranking compares nDCG@10 and essential passages instead of passage-level errors, which change under any reordering.

Table 5. Exit status of run and score.
StatusMeaning
0Completed; the gate passed or no gate was requested.
2Completed; the gate failed.
1Contract or execution error; nothing is scored.

3.3Offline rescoring

score --run DIR rescores a completed run from saved outputs without calling a model. It can re-pair with other baselines or controls, change the gate or targets, and re-key the run against a corrected revision of the same dataset when every input is unchanged, so an answer-key correction does not need a paid rerun.

3.4Review

Each run writes review-brief.json, which groups changed cases into worse, better, noisy and remaining, by category. review draft turns the brief into an unexamined draft with one numbered finding per group. A reviewer edits it into interpretation.json with evidence for every examined case. A later run can target the cases behind a recommendation with --from-review RUN_DIR --recommendation N. Reviews never change scores or the gate.

3.5Report

report.html is one file that makes no network requests: the verdict, headline metrics with changes from the baseline, coverage, and every case grouped by bucket and category, with a viewer for the type (table grid, image overlay or JSON). Image viewers load case images from a sibling report-assets/ directory. serve --run DIR serves a run read-only on 127.0.0.1 for browsers that block local file links.

4Datasets

A dataset is a directory with manifest.json, cases.jsonl, a separate answers.jsonl and the source files the cases use, each pinned by size and SHA-256. Its identity (ID, version and digest) is recorded in every run, and a baseline from another revision is refused. A revision may retire cases with a stated reason.

Table 6. Kinds of dataset and where their files are stored.
KindContentsLocationStatus
Synthetic smoke Generated pages, crops, tables and queries that exercise each type's scoring and report. Core repository. Image files may also be served from a public object store over HTTPS without credentials. Available in 0.2.0
Private suites Cases derived from production documents and human review, held by the organization that owns them. A private plugin repository; source files in a private object store. Used for regression testing; not published
Public development set A fictional site file written from licensed public information: maps, tables, lab results and reports, with answers. Core repository Planned
Sealed holdout A larger set in the same formats. Version, size, selection method and checksum published; answers withheld. Maintainer Planned

Core refuses to upload a private dataset to a public store, or to point a public dataset at a private one. Uploads verify each file's checksum, never replace a stored object with different bytes and never delete. Promoted baseline and control runs can be archived to a private store and are restored, checksum-verified, before a paired run on another machine.

5Running a benchmark

Requirements: Ruby 3.3.2. There are no third-party gems.

bin/test
bin/envirobench benchmarks
bin/envirobench datasets
bin/envirobench validate-dataset benchmarks/table-separators/datasets/synthetic-v1

Run a synthetic dataset through the reference adapter, which calls no model, then compare a second run with the first and open the report:

bin/envirobench run --adapter examples/reference-adapter --run separators-base \
  --model reference --benchmark table-separators \
  --dataset table-separators=envirobench/table-separators-synthetic-v1 \
  --adapter-option variant=baseline --budget 0
bin/envirobench run --adapter examples/reference-adapter --run separators-new \
  --model reference --benchmark table-separators \
  --dataset table-separators=envirobench/table-separators-synthetic-v1 \
  --baseline-run results/separators-base --gate logical --require-no-regressions --budget 0
bin/envirobench serve --run results/separators-new --open

Each run writes one directory:

results/<run-id>/
  run.json             how the run was made, dataset identity, gate result
  adapter-request.json input-only request
  adapter/             adapter-owned files, never modified by core
  comparison.json      answers joined with every compared run's predictions
  metrics.json         per-profile scores, gates and per-case errors
  report.html          the report; works offline
  report-assets/       case images for image viewers
  summary.md           plain-English verdict and next steps
  review-brief.json    changed cases grouped for a reviewer

5.1Saved profiles

A plugin can save how it runs each benchmark in benchmarks/NAME/bench.json: the dataset, adapter and options, gate, budget, the promoted baseline and controls per adapter, and a run store. bench turns a profile into the full run command:

bin/envirobench bench run table-separators --plugin ../acme-plugin --unpaired --run base
bin/envirobench bench promote table-separators ../acme-plugin/.private/runs/base \
  --control ../acme-plugin/.private/runs/control --plugin ../acme-plugin
bin/envirobench bench run table-separators --plugin ../acme-plugin --run my-change --plan
bin/envirobench bench run table-separators --plugin ../acme-plugin --run my-change

--adapter NAME|PATH selects the system under test. Baselines are kept per adapter, so a new adapter runs --unpaired until a baseline is promoted for it. --plan prints the commands without writing files or calling a model.

6Extending EnviroBench

6.1Adding a benchmark type

A new type is one Ruby file, a synthetic dataset and two prediction fixtures:

lib/envirobench/benchmarks/table_separators.rb       the type
benchmarks/table-separators/datasets/synthetic-v1/   manifest.json, cases.jsonl, answers.jsonl, assets/
test/fixtures/benchmarks/table-separators/           correct.json and wrong.json
envirobench-plugin.json                              one dataset entry

The type declares ID, FORMAT, LABEL and SCORING_VERSION, optionally PROFILES, HEADLINE, CATEGORIES and a viewer, and implements score_case. Core supplies pairing, the gate, rescoring, review and the report. bin/test runs a shared conformance suite over every registered type: answer isolation, malformed results, correct and wrong fixtures, pairing and gates, re-keying, review drafts and the report viewer.

6.2Writing an adapter

An adapter is any executable that reads the request and writes adapter-result.json to its output directory, naming the adapter, its version, the spend and a predictions file with exactly the requested case IDs:

{"id": "table-separators", "status": "completed",
 "format": "envirobench.table-separators.v1", "data": "predictions.json",
 "systems": [{"role": "candidate", "model": "provider/model", "cost_usd": 0.42,
              "execution_failures": 0, "unpriced_attempts": 0}]}

Unresolved execution failures, unpriced model calls and missing, duplicate or malformed results block scoring. An adapter may save logs and traces; EnviroBench does not score them. examples/reference-adapter implements the protocol for every public synthetic dataset.

7Status and limitations

  • Adapters are trusted local programs. Third-party adapters cannot yet be run with limited file, network and process access.
  • There are no repeated trials, confidence intervals or LLM judges. Same-code control runs are the current check for noise.
  • The included datasets are small. The largest, 40 invented documents for document extraction, is generated with its answers, so its tables follow a limited set of layouts. The public development set and sealed holdout do not exist yet.
  • EnviroBench is in development and its test sets are being finalized. No scores are published yet; they will be at release, first on the public synthetic test sets, then on the planned development set and sealed holdout.
  • The repository, package and dataset licenses have not been chosen.

8Versions

Table 7. Changes to the core.
DateChange
October 2026 One benchmark contract for every type. New types: table-separators, table-containers, ocr-auto-accept, search-reranking. Repeatable control runs. Saved profiles (bench) with adapter selection, dataset revisions, and public and private object stores. Category scores (Extract, Verify, Locate, Research) built from per-task results. Document-style synthetic cases: a drafted site plan, a typewritten report page and table, and a laboratory summary table. New type document-extraction (a whole PDF to measurement records) with shared record matching and normalisers, and 40 synthetic documents in nine kinds; category scoring version 2 adds it to Extract.
September 2026 table-mapping with exact and projected scores, offline re-keying, review briefs and drafts, and a loopback report server.
August 2026 Version 0.2.0: benchmark plugins with checksummed assets, and leaderboards across saved runs.
August 2026 Product-neutral core with sample-location-placement and table-cell-ocr.

9Citation

Until a paper is available, cite the software:

@software{envirobench,
  title   = {EnviroBench: a benchmark for AI systems that read environmental documents},
  author  = {{Statvis}},
  year    = {2026},
  version = {0.2.0},
  url     = {https://envirobench.com},
  note    = {Pre-release}
}

Questions: hello@statvis.com. For an overview of the tasks in environmental terms, see the home page.