EnviroBench
Methodology: tasks, metrics, evaluation protocol and running the benchmark
- Version
- 0.2.0 (pre-release)
- Benchmark types
- 17
- Runtime
- Ruby 3.3, standard library only
- Repository
- Private until release
EnviroBench is not yet public. The code runs today on small synthetic datasets; the public development set, sealed holdout and license are still in preparation (§7). This page describes version 0.2.0 as implemented and marks planned work as planned.
Abstract
EnviroBench tests AI systems on document tasks that environmental teams do: reading analytical table cells, mapping table cells to samples and measurements, locating tables and their row and column lines on scanned pages, placing sample locations on site plans, deciding which OCR readings can skip human review, and ranking passages for search. Each benchmark type fixes an input format, an answer format and deterministic per-case scoring. The system under test runs behind an adapter that receives inputs only. EnviroBench joins the held-back answers afterwards, records every error on every case, and compares a run case by case with a saved baseline and same-code control runs. A change passes only when it introduces no new error on a case that is stable across repeats; a better average never cancels a new error.
1Design
EnviroBench separates what is measured from what is being measured. The public core defines benchmark types, answer formats, scoring rules, the adapter protocol and the report. It does not run any product or model. Datasets come from plugins, and each system under test is an adapter. An organization can keep its own plugin, adapter and object store private while using the same public scoring.
| Component | Owns | Never does |
|---|---|---|
| Core | Benchmark types (one Ruby file each), the run lifecycle, scoring, pairing, gates, review drafts and the HTML report. Public synthetic datasets. | Run a product, call a model, or hold production-derived data. |
| Plugin | Datasets: cases, answer keys, checksummed source files, optional source links, and saved run profiles (bench.json). |
Define scoring. Plugins add data only. |
| Adapter | Runs one system or model on input-only cases and returns its predictions, cost and usage. | See answers, earlier runs or source links, or score itself. |
2Tasks and metrics
A benchmark type is one class in lib/envirobench/benchmarks/. It declares an
identifier, a result format, a scoring version and its score profiles, and implements
score_case, which returns, for each profile, the case's errors and metrics. Errors
are comparable values, so the same mistake in two runs is recognized as the same error.
| Benchmark | Input | Prediction | Unit | Synthetic |
|---|---|---|---|---|
document-extraction |
One whole PDF (digital or scanned) and a statement of which values are in scope | Every field-sample measurement as a record: location, sample, date, depth, matrix, parameter, result, unit, non-detect and limit, qualifiers | document | 40 |
site-dataset |
A site's whole document record: every PDF, most pages holding no results | The site's deduplicated analytical dataset, each record citing a page that shows its governing value | site | 1 |
site-questions |
The same record and a fixed question catalogue | Typed, cited answers, or "not in the record" | site | 1 |
sample-location-placement |
Site plan image and a sample-location label | Whether the marker was found, and its point in percent of the image | point | 24 |
table-cell-ocr |
Crop of one table cell and the existing first-pass reading | The cell's text | cell | 80 |
table-mapping |
A table as CSV, with optional cell boxes | Orientation, a role for each cell (sample identity, chemical name, value, unit, method), value groups and default units | table | 37 |
table-separators |
Table crop image, with optional word boxes | Positions of vertical and horizontal separator lines | crop | 13 |
table-containers |
Page image, with optional detector boxes and word boxes | One box per table, and an optional angle that straightens the page | page | 11 |
ocr-auto-accept |
Saved readings of one cell from several OCR and vision readers, plus image checks | Accept one value, or send the cell to review | cell | 100 |
search-reranking |
A query and 1–100 retrieved passages | An ordering of passage IDs | query | 46 |
instrument-execution |
Page images of a legal instrument and any related parent, schedule or reference pages, with its title, type and the parties expected to sign | Execution status (executed, partial, unsigned, unknown, missing, referenced only), the parties visibly signed and the execution page | instrument | 51 |
party-roles |
Page images of one document (tank registration, lease, consignment agreement, title, permit, notice or letter), its title and type, and the parties it names | The roles the document states for each party: owner, tank owner, operator, registrant, lessor, lessee, consignor, consignee or other | document | 42 |
indemnity-terms |
A contract's definitions, indemnity and release clauses, limitations, survival, schedules and amendments as page images, with its title, type and parties | Each indemnity and release: kind, giver and beneficiary, environmental coverage, period before or after closing, cap, survival and carve-outs | contract | 41 |
chain-of-title |
A site's title documents (certificates, searches, abstracts) as page images, including a neighbouring parcel's search, and the subject parcel's legal description | The subject parcel's ownership periods: owners, start and end dates as printed, granting instrument, consideration, and whether a period continues the same owner under a new name | chain | 40 |
rosc-fields |
An AER Record of Site Condition form (OneStop Submission PDF); a public real-world set of 84 forms downloads from AER at pinned URLs | Submission ID and date, licensee and its identifier, site name, intents, and the UWIs and related submissions listed, each with a quote | form | 4 |
source-conflicts |
A site record of 7–10 documents (assessments, laboratory, regulator, closure, tank, title and spill records) as page images, with a document list | Each material conflict: fact type, both sides cited by document and page, and the governing source where a rule settles it | site record | 30 |
missing-information |
A site record with documents withheld, and questions from a fixed catalogue (owner, tanks, last groundwater event, remediation, regulatory status, spills, latest exceedance) | Each answer with a citation, or not in the record; the documents the record cites but does not hold | site record | 30 |
2.1Errors and headline metrics
Each profile is a separate way of scoring the same prediction. The first profile is the default gate. Headline metrics summarize a run; the gate uses only per-case errors.
| Benchmark | Profile | A case has an error when | Headline metrics |
|---|---|---|---|
document-extraction |
records | An expected record has no predicted match, or a predicted record matches nothing. Records match on parameter (CAS number, a public synonym table, or the name) and sample identity, then by an optimal one-to-one assignment. | Correct-record F1, pooled over all documents: 2 × matched records whose value is right (non-detect state, number or limit, unit, and reporting limit where the answer gives one) ÷ (expected + predicted records). Record F1, counting every matched record whatever its value. |
| values | On a matched record, the result (after unit conversion), non-detect state, unit, reporting limit or qualifiers differ. A wrong value with no non-detect, qualifier or review flag is also a silent wrong value. | Value accuracy (non-detect state, number and unit right); silent wrong values as a share of matched records, lower is better. | |
| metadata | On a matched record, the location, sample ID, date, time, depth or matrix differ. | Field accuracy per field. | |
site-dataset |
records | A measurement is missing, invented, or reported again (a duplicate of one already reported). | Correct-record F1 (right governing value and a valid citation); duplicate rate, lower is better; record F1. |
| values | As document extraction, plus a superseded value: a number a correction letter or a final report replaced. | Silent wrong values, lower is better. | |
| metadata | As document extraction, plus a citation to a page that does not show the governing value. | Valid citations. | |
site-questions |
default | An answer is wrong or missed, an answer is given where the record has none (invented, severe), or a right answer cites no supporting page (severe). | Questions right; invented answers and unsupported citations, lower is better; data questions right. |
sample-location-placement |
default | The point is more than 1% of the image diagonal from the checked location: near (within 5%), far, outside the image, or the marker was not found. | Points within 1% of the diagonal; within 5%; labels found; median error as % of diagonal. |
table-cell-ocr |
default | The reading differs from the checked text after whitespace normalization; a wrong reading that equals the first-pass value is also a false accept, because it would skip review. | Cells read exactly; wrong auto-accepts as a share of auto-accepted readings; share auto-accepted. |
table-mapping |
exact | Any orientation, cell role, value group or default unit is missing or extra. | Tables matching the expected mapping exactly. |
| important | An important fact differs after a plugin-supplied projector turns the mapping into records (for example, the measurements a product would import). | Available important checks; measurement correctness; sample assignment. | |
table-separators |
strict | An expected line has no predicted line within 5 units (matched nearest-first, one to one), or a predicted line matches nothing. | Line F1 pooled over all crops and both axes; crops with every line right; line edits per crop. |
| logical | As strict, after replacing each line with the number of word centres before it, so lines that split the same words are equal. | Crops splitting the words correctly. | |
| word pairs | Words are grouped into different rows or columns than in the expected grid. | Word-pair F1, rows and columns averaged. | |
table-containers |
boxes | An expected box has no predicted box at IoU ≥ 0.3, a predicted box matches nothing, or a matched box has an edge more than 10 units off. | Pages needing no box edit; mean IoU; missed and extra boxes per page. |
| angle | The straightening angle is more than 0.3° from the expected angle. | Pages straightened within 0.3°. | |
ocr-auto-accept |
default | An accepted value differs from the checked value, the cell was not fully legible, or blank and non-blank are confused. Sending a cell to review is never an error. | Legible cells auto-accepted correctly; wrong auto-accepts; cells sent to review per page. |
search-reranking |
default | A judged relevant passage is outside the top 10, or a judged irrelevant passage is inside it. For the gate, a query regresses when nDCG@10 falls by more than 0.01 or fewer essential passages stay in the top 10. | nDCG@10 (graded gain, log2 discount); relevant passages in the top 10; queries keeping every essential passage. |
instrument-execution |
default | The status differs (an instrument called executed when not fully executed, and an unreadable, absent or only-referenced execution page called unsigned, are their own errors), a signed party is missed or an unsigned one listed, or the execution page differs. | Statuses correct; instruments called executed when not; evidence gaps called unsigned; signatory recall and precision. |
party-roles |
default | A party is given a role the document does not state for it (an owner or tank-owner role is its own error) or is missing a stated role. | Parties with every role right; non-owners called owners; stated roles found; predicted roles the document states. |
indemnity-terms |
default | A provision is missing or invented, or its environmental coverage, period, cap, survival or carve-outs differ; environmental coverage claimed where the contract does not give it, and a dropped environmental carve-out, are their own errors. | Contracts read exactly right; contracts overstating environmental protection; provisions found; provision fields right; carve-outs found. |
chain-of-title |
default | A period is missing, or its instrument, dates, consideration or continuation differ; a period from the neighbouring parcel, and an owner the parcel never had, are their own errors. | Periods exactly right; periods from another parcel or invented; owners found; dates right as printed; chains exactly right. |
rosc-fields |
default | A field's value differs from the form (identifiers compared without spacing), a value is missed, a list item is missed or added, or a blank field is filled in. | Fields read right; forms read exactly right; blank fields filled in; values with a matching quote. |
source-conflicts |
default | A conflict is missed or cited on the wrong page, or a non-conflict is reported (restatements, replaced drafts, other sampling events and a neighbour's facts are their own kinds); fact type and governing source are also checked. | Conflicts found and cited; conflicts found; false conflicts; both sides cited right. |
missing-information |
default | An answer is wrong, missed or unsupported by its citations, a question the record cannot answer is answered (an invented answer), or an absent document is missed or misreported. | Questions right; invented answers; answerable questions right; absent documents found. |
Scoring versions are recorded in every run, for example
envirobench.table-separators.scores.v1. A change to a formula gets a new version.
2.2Tiers
Each benchmark type declares a tier: what a case hands the system. The component tier hands it a crop, a cell or a page; the document tier one whole PDF; the site tier a site's whole document record, a few hundred pages of reports, certificates, letters, invoices and scans, most of them holding nothing that is asked for. Real work starts at the site tier: finding the pages that matter, reading them, and reconciling the same results reprinted across years of reports, drafts and finals, scanned copies and correction letters. The lower tiers show where a system loses accuracy on the way.
At the site tier cost and wall time per site are part of the result, reported beside every score: a system that is accurate only at a high price per site is a different product from one that is accurate cheaply, so site boards sort by both. Three public sites (a service station, a bulk plant on railway lands and a dry cleaner, from 275 to 806 pages) are invented and generated with their answers; real site records stay in private plugins.
2.3Category scores
Leaderboards summarise the tasks in six categories: Extract, Verify, Locate, Research, Legal and Site. Each component task is first turned into a score from 0 to 100, higher better. By default the component score is 100 × the task's primary headline measure; every current primary measure is a share or a mean between 0 and 1 where higher is better, and a lower-is-better primary would need its own published formula. Two tasks built around one dangerous error are the exceptions. OCR auto-accept scores 100 × max(0, (correct accepts − 10 × wrong accepts) ÷ readable cells), so one wrong value that would skip review cancels ten correct accepts and coverage cannot buy back a wrong accept cheaply. Instrument execution scores 100 × max(0, (correct statuses − 5 × instruments called executed when not) ÷ instruments), so a release or transfer wrongly reported as signed cancels five correct statuses, and calling every instrument executed scores 0. Party roles scores 100 × max(0, (parties with every role right − 5 × non-owners called owners) ÷ parties), the same weight for a party wrongly made an owner, and indemnity terms 100 × max(0, (contracts read exactly right − 5 × contracts overstating environmental protection) ÷ contracts), and chain of title 100 × max(0, (periods exactly right − 5 × periods from another parcel or invented) ÷ periods). In Research, missing information scores 100 × max(0, (questions right − 5 × invented answers) ÷ questions).
A category score is the weighted mean of its component scores, Σ wisi ÷ Σ wi. Weights are per task, not per case, so a large test set does not outweigh a small one. Every weight is 1 except document extraction in Extract.
Document extraction is Extract's end-to-end task: a whole PDF in, measurement records out. It is the outcome environmental professionals actually use, so it has weight 4, as much as the four table tasks together (4 of the category's 9). The table tasks show where a system loses records on the way. Its primary measure is correct-record F1: a record counts only when it is found and its value is right.
- Complete or nothing. A system gets a category score only when it has results for every component task, on the test-set versions the leaderboard shares. Otherwise the category shows Incomplete (n of m tasks). A missing task is never counted as zero and never dropped.
- Repeats. With repeated runs, the score is the weighted mean of each task's mean, shown with a range: the weighted mean of each task's lowest and highest repeat.
- No overall score. The categories cover very different amounts of work, so they are not combined into one number.
The definition lives in core (lib/envirobench/categories.rb) with its own scoring
version; the table below is generated from it.
| Category | Component task | Tier | Weight | Component score (0–100) |
|---|---|---|---|---|
| EnviroBench ExtractHow accurately does the system turn documents into data? | document-extraction | document | 4 | 100 × correct-record F1 |
table-containers | component | 1 | 100 × pages needing no box edit | |
table-separators | component | 1 | 100 × line F1 (all lines) | |
table-cell-ocr | component | 1 | 100 × cells read exactly | |
table-mapping | component | 1 | 100 × exact mapping agreement | |
rosc-fields | document | 1 | 100 × fields read right | |
| EnviroBench VerifyDoes the system know when it might be wrong? | ocr-auto-accept | component | 1 | 100 × max(0, (correct accepts − 10 × wrong accepts) ÷ readable cells) |
| EnviroBench LocateCan the system place things on site plans? | sample-location-placement | component | 1 | 100 × within 1% of the image diagonal |
| EnviroBench ResearchDoes the system find the right evidence, and see where the record conflicts or falls short? | search-reranking | component | 1 | 100 × nDCG@10 |
source-conflicts | site | 1 | 100 × conflicts found and cited | |
missing-information | site | 1 | 100 × max(0, (questions right − 5 × invented answers) ÷ questions) | |
| EnviroBench LegalDoes the system read legal instruments without overstating them? | instrument-execution | document | 1 | 100 × max(0, (correct statuses − 5 × called executed when not) ÷ instruments) |
party-roles | document | 1 | 100 × max(0, (parties with every role right − 5 × non-owners called owners) ÷ parties) | |
indemnity-terms | document | 1 | 100 × max(0, (contracts read exactly right − 5 × contracts overstating environmental protection) ÷ contracts) | |
chain-of-title | site | 1 | 100 × max(0, (periods exactly right − 5 × periods from another parcel or invented) ÷ periods) | |
| EnviroBench SiteCan the system work from a site's whole document record? | site-dataset | site | 1 | 100 × correct-record F1 |
site-questions | site | 1 | 100 × questions right |
3Evaluation protocol
One run covers one benchmark and one system. Figure 1 shows the lifecycle. The boundary between core and adapter is the only place data crosses, and answers never cross it.
- CoreLoad dataset. Cases, answers and checksummed source files.
- CoreWrite an input-only request. No answers, earlier runs or source links.
- AdapterRun the system. Return candidate predictions and cost.
- CoreValidate and join. Add answers, the baseline and controls.
- CoreScore every case. Errors and metrics per profile.
- CorePair case by case. New errors; unstable cases masked.
- CoreGate and report. Pass, fail or baseline required.
3.1Answer isolation
The adapter is started as ADAPTER run --request FILE --output DIR. The request
holds the run ID, model, spend limit, adapter options and an input-only view of the dataset:
cases and a run-specific directory of checked source files. It never contains answer rows, the
answer file's path, earlier runs' predictions or source links. The adapter returns only the
candidate's predictions. Core checks each one with the type's validator; a technical failure
blocks the report (exit status 1) and is never counted as a score.
3.2Paired comparison and the regression gate
A run is compared with saved runs rather than with a second model in the same run:
--baseline-run DIRis the run the change is compared with.--control-run DIR, repeatable, is a same-code repeat of the baseline. A case is unstable in a profile when the baseline and any control disagree. Its regressions are reported but do not fail the gate. More controls catch more noise.--target-case IDnames cases the change must get right. Any remaining error on a target fails the gate.
Per profile, a case regresses when the candidate has an error the baseline did not have. The gate passes when no stable case regressed and every target is error-free. Without a baseline the status is baseline required; when a profile cannot score every case and run, it is unavailable. A type may redefine when a case is worse: search reranking compares nDCG@10 and essential passages instead of passage-level errors, which change under any reordering.
| Status | Meaning |
|---|---|
| 0 | Completed; the gate passed or no gate was requested. |
| 2 | Completed; the gate failed. |
| 1 | Contract or execution error; nothing is scored. |
3.3Offline rescoring
score --run DIR rescores a completed run from saved outputs without calling a
model. It can re-pair with other baselines or controls, change the gate or targets, and re-key
the run against a corrected revision of the same dataset when every input is unchanged, so an
answer-key correction does not need a paid rerun.
3.4Review
Each run writes review-brief.json, which groups changed cases into worse, better,
noisy and remaining, by category. review draft turns the brief into an unexamined
draft with one numbered finding per group. A reviewer edits it into
interpretation.json with evidence for every examined case. A later run can target
the cases behind a recommendation with --from-review RUN_DIR --recommendation N.
Reviews never change scores or the gate.
3.5Report
report.html is one file that makes no network requests: the verdict, headline
metrics with changes from the baseline, coverage, and every case grouped by bucket and category,
with a viewer for the type (table grid, image overlay or JSON). Image viewers load case images
from a sibling report-assets/ directory. serve --run DIR serves a run
read-only on 127.0.0.1 for browsers that block local file links.
4Datasets
A dataset is a directory with manifest.json, cases.jsonl, a separate
answers.jsonl and the source files the cases use, each pinned by size and SHA-256.
Its identity (ID, version and digest) is recorded in every run, and a baseline from another
revision is refused. A revision may retire cases with a stated reason.
| Kind | Contents | Location | Status |
|---|---|---|---|
| Synthetic smoke | Generated pages, crops, tables and queries that exercise each type's scoring and report. | Core repository. Image files may also be served from a public object store over HTTPS without credentials. | Available in 0.2.0 |
| Private suites | Cases derived from production documents and human review, held by the organization that owns them. | A private plugin repository; source files in a private object store. | Used for regression testing; not published |
| Public development set | A fictional site file written from licensed public information: maps, tables, lab results and reports, with answers. | Core repository | Planned |
| Sealed holdout | A larger set in the same formats. Version, size, selection method and checksum published; answers withheld. | Maintainer | Planned |
Core refuses to upload a private dataset to a public store, or to point a public dataset at a private one. Uploads verify each file's checksum, never replace a stored object with different bytes and never delete. Promoted baseline and control runs can be archived to a private store and are restored, checksum-verified, before a paired run on another machine.
5Running a benchmark
Requirements: Ruby 3.3.2. There are no third-party gems.
bin/test
bin/envirobench benchmarks
bin/envirobench datasets
bin/envirobench validate-dataset benchmarks/table-separators/datasets/synthetic-v1
Run a synthetic dataset through the reference adapter, which calls no model, then compare a second run with the first and open the report:
bin/envirobench run --adapter examples/reference-adapter --run separators-base \
--model reference --benchmark table-separators \
--dataset table-separators=envirobench/table-separators-synthetic-v1 \
--adapter-option variant=baseline --budget 0
bin/envirobench run --adapter examples/reference-adapter --run separators-new \
--model reference --benchmark table-separators \
--dataset table-separators=envirobench/table-separators-synthetic-v1 \
--baseline-run results/separators-base --gate logical --require-no-regressions --budget 0
bin/envirobench serve --run results/separators-new --open
Each run writes one directory:
results/<run-id>/
run.json how the run was made, dataset identity, gate result
adapter-request.json input-only request
adapter/ adapter-owned files, never modified by core
comparison.json answers joined with every compared run's predictions
metrics.json per-profile scores, gates and per-case errors
report.html the report; works offline
report-assets/ case images for image viewers
summary.md plain-English verdict and next steps
review-brief.json changed cases grouped for a reviewer
5.1Saved profiles
A plugin can save how it runs each benchmark in benchmarks/NAME/bench.json: the
dataset, adapter and options, gate, budget, the promoted baseline and controls per adapter, and
a run store. bench turns a profile into the full run command:
bin/envirobench bench run table-separators --plugin ../acme-plugin --unpaired --run base
bin/envirobench bench promote table-separators ../acme-plugin/.private/runs/base \
--control ../acme-plugin/.private/runs/control --plugin ../acme-plugin
bin/envirobench bench run table-separators --plugin ../acme-plugin --run my-change --plan
bin/envirobench bench run table-separators --plugin ../acme-plugin --run my-change
--adapter NAME|PATH selects the system under test. Baselines are kept per adapter,
so a new adapter runs --unpaired until a baseline is promoted for it.
--plan prints the commands without writing files or calling a model.
6Extending EnviroBench
6.1Adding a benchmark type
A new type is one Ruby file, a synthetic dataset and two prediction fixtures:
lib/envirobench/benchmarks/table_separators.rb the type
benchmarks/table-separators/datasets/synthetic-v1/ manifest.json, cases.jsonl, answers.jsonl, assets/
test/fixtures/benchmarks/table-separators/ correct.json and wrong.json
envirobench-plugin.json one dataset entry
The type declares ID, FORMAT, LABEL and
SCORING_VERSION, optionally PROFILES, HEADLINE,
CATEGORIES and a viewer, and implements score_case. Core supplies
pairing, the gate, rescoring, review and the report. bin/test runs a shared
conformance suite over every registered type: answer isolation, malformed results, correct and
wrong fixtures, pairing and gates, re-keying, review drafts and the report viewer.
6.2Writing an adapter
An adapter is any executable that reads the request and writes
adapter-result.json to its output directory, naming the adapter, its version, the
spend and a predictions file with exactly the requested case IDs:
{"id": "table-separators", "status": "completed",
"format": "envirobench.table-separators.v1", "data": "predictions.json",
"systems": [{"role": "candidate", "model": "provider/model", "cost_usd": 0.42,
"execution_failures": 0, "unpriced_attempts": 0}]}
Unresolved execution failures, unpriced model calls and missing, duplicate or malformed results
block scoring. An adapter may save logs and traces; EnviroBench does not score them.
examples/reference-adapter implements the protocol for every public synthetic
dataset.
7Status and limitations
- Adapters are trusted local programs. Third-party adapters cannot yet be run with limited file, network and process access.
- There are no repeated trials, confidence intervals or LLM judges. Same-code control runs are the current check for noise.
- The included datasets are small. The largest, 40 invented documents for document extraction, is generated with its answers, so its tables follow a limited set of layouts. The public development set and sealed holdout do not exist yet.
- EnviroBench is in development and its test sets are being finalized. No scores are published yet; they will be at release, first on the public synthetic test sets, then on the planned development set and sealed holdout.
- The repository, package and dataset licenses have not been chosen.
8Versions
| Date | Change |
|---|---|
| October 2026 | One benchmark contract for every type. New types: table-separators,
table-containers, ocr-auto-accept, search-reranking.
Repeatable control runs. Saved profiles (bench) with adapter selection,
dataset revisions, and public and private object stores. Category scores (Extract, Verify, Locate, Research) built from per-task results. Document-style synthetic
cases: a drafted site plan, a typewritten report page and table, and a laboratory
summary table. New type document-extraction (a whole PDF to measurement
records) with shared record matching and normalisers, and 40 synthetic documents in nine
kinds; category scoring version 2 adds it to Extract. |
| September 2026 | table-mapping with exact and projected scores, offline re-keying, review
briefs and drafts, and a loopback report server. |
| August 2026 | Version 0.2.0: benchmark plugins with checksummed assets, and leaderboards across saved runs. |
| August 2026 | Product-neutral core with sample-location-placement and
table-cell-ocr. |
9Citation
Until a paper is available, cite the software:
@software{envirobench,
title = {EnviroBench: a benchmark for AI systems that read environmental documents},
author = {{Statvis}},
year = {2026},
version = {0.2.0},
url = {https://envirobench.com},
note = {Pre-release}
}
Questions: hello@statvis.com. For an overview of the tasks in environmental terms, see the home page.