GroundedThe argument

Research programme · draft, not submitted

A wrong answer that survives review

Asked about a rupture disk, an assistant answered about a relief valve. Nothing signalled the swap. The public safety data meant to catch it is no more trustworthy.

Idle
The failure, reconstructed
An engineer asks about sizing overpressure protection for a reactor protected by a rupture disk. The assistant answers with content about a spring-loaded pressure relief valve: set pressure, blowdown, reseating. None of these quantities exist for a rupture disk, which is a one-shot non-reclosing device.
Wording illustrative · parameters from the benchmark item

The pivot to data

Raw vs screened national TRIR, 2019 249x

One reporting year. One plausibility screen. No single multiplier repairs a benchmark built on raw OSHA filings.

Flagged again next year 27.2%

Versus 0.95% for an establishment never flagged before.

Built in response 6 repos 2 papers

Data quality, a grounding benchmark, an ontology, a simulation, a calculator, a paper trail. No language model has been scored on the benchmark yet.

Safety numbers are only as good as the denominator and the citation underneath them. Programme status: see changelog for current counts.
00

Five moves run from a wrong answer to one question you can ask on Monday

Each move states its finding and the number behind it. The figure on the left changes with the move.

Source: ehs-ai-grounding-eval, corpus/items/pressure_relief.json
Move 01

Zero parameters are shared by the device asked about and the device answered about

A rupture disk question came back with set pressure, blowdown and reseating. Those are relief valve parameters; a rupture disk has none of them. Nothing in the answer contradicts itself, so a consistency check finds no tell.

0 shared parameters
Move 02

One plausibility screen on hours moves the 2019 national rate 249 times

In reporting year 2019 the raw national TRIR is 0.016 and the screened one is 4.018, a ratio of 249.3. Across 2,801,064 real OSHA filings from 2016 to 2024 the raw series swings from 0.016 to 2.92 while the screened series stays between 3.15 and 4.60.

249x, 2019
Move 03

A flagged site is flagged again 27.2 percent of the time, against 0.95 percent for the rest

That is an odds ratio of 39.1. A permutation null centred at 1.0 rules out independent keying errors as the main mechanism. The bad hours sit with the same filers year after year, so the noise does not average out.

39.1odds ratio
Move 04

Six repositories answer one failure each, in three lanes

One is empirical, two are apparatus, one is simulation, two are tools. They bind the answer to a source, test the denominator under the rate, and price what the decision costs. Each one ships its own data and its own story page.

6repositories
Move 05

Three of the fifteen headline claims have not been tested at all

Nine of the fifteen are verified against files in the repositories and three are drafts. No language model has been run through the benchmark, the SEM work is entirely synthetic, and both papers are unsubmitted. All of that is in the ledger below, not in a footnote.

9 / 3 / 3verified / draft / untested
Monday

One question separates a rate from a number: what is the hours screen?

On this data the screen removes 96.68 percent of reported hours. Ask any vendor or internal report for its hours-per-employee screen and for the share of hours that screen removes. If there is no screen, the rate is not a rate.

1question

A rupture disk question came back with relief valve parameters, and it read as competent

Two parameter spaces that do not intersect, and nothing in the wording signals the swap.

01

Asked about a rupture disk, answered about a relief valve

Fluent, well formatted, and inapplicable to the device in front of the engineer. The two devices share no parameters, so every value in the answer is plausible for the device it was not about.

Asked about

Rupture disk

Opens once and stays open.

  • Burst pressure
  • Burst tolerance
  • Manufacturing design range
  • Operating ratio
shared parameters
Came back

Relief valve

Parameters of a device that recloses. A rupture disk has none of them.

  • Set pressure
  • Blowdown
  • Reseating
  • Chatter
Why the substitution survives review

Nothing in the answer is wrong on its own terms

The two parameter spaces do not intersect, so every value is a plausible value for a relief valve. A reviewer checking internal consistency finds none of the usual tells: no contradictions, no missing units, no hedging. The error lives in the binding between the question and the device, and that binding was never made.

Authoritative binding replaces free retrieval: every claim carries a source clause and an anchor span. A benchmark scores abstention as a correct response when the corpus does not contain the answer. The ontology makes the device and context explicit before any parameter is named. And the data work underneath asks whether the public benchmarks used to check any of this are themselves trustworthy.

Why it recurs: four mechanisms converge
  1. 1

    No retrieval happened

    The answer came from parametric memory, not a bound source.

  2. 2

    Correct content is long tail

    Relief valves dominate the training distribution. Disks do not.

  3. 3

    Abstention is penalised

    A confident answer scores better than "not in my sources".

  4. 4

    The corpus is lexically adversarial

    Standards share vocabulary across devices that share no parameters.

A plausibility screen on hours moves the 2019 national rate by a factor of 249

Any benchmark quoted off raw OSHA filings can be wrong by two orders of magnitude.

02

Screening hours turns a series that swings 0.016 to 2.92 into one that holds near four

The screen is a plausibility test on hours per employee. It flags 2.07% of filings, and those filings carry 96.68% of all reported hours.

Filings analysed2,801,064

Real public OSHA ITA data, 2016 to 2024.

Filings flagged2.07%

carry 96.68% of all reported hours.

Pooled TRIR0.134 3.983

Unscreened to screened, all years pooled.

Pooled ratio29.7x

Median screened establishment TRIR is 2.47.

Aggregate TRIR by reporting year. Screening produces a stable series. The raw file does not.

Unscreened TRIRScreened TRIRHover, tap, or focus a year and use arrow keys

A filing is implausible when reported hours per employee fall outside 120 to 4,500. Pooled across nine years that flags 2.07% of 2,801,064 filings, which carry 96.68% of reported hours. The pooled ratio of screened to unscreened TRIR is 29.7x; the median screened establishment TRIR is 2.47. The bounds are a defensible choice rather than a discovered truth, which is why a sensitivity grid ships with the repository. The analysis measures what employers filed, not workplaces.

Source: ehs-osha-analysis, outputs/summary.json (quality.by_year) and outputs/tables/quality_by_year.csv

The same filers keep getting flagged. Probability of a flag this year, by last year's status.

Odds ratio 39.1
Flagged last year27.2%

5,776 of 21,243 establishments flagged again.

Not flagged last year0.95%

13,394 of 1,415,236 establishments newly flagged.

Pooled

Each grid is 100 establishments. A site flagged for implausible hours one year is 39.1 times more likely to be flagged again the next. A permutation null centred at 1.0 (95% range 0.92 to 1.09) rules out independent year-to-year keying errors as the dominant mechanism.

Source: ehs-osha-analysis, outputs/tables/panel_transitions.csv and docs/PANEL.md

Open the Observatory: every chart, every source file

The trap is a clause that reads correct for the wrong device

Three real items from the 68-item corpus. Answer them before you read how the baselines did.

03

TF-IDF retrieval scores 27.0 percent, below a 52.1 percent random floor

Pick an answer on three real items from the 68-item corpus, then see the co-hyponym trap and how the three non-LLM baselines scored the same item. Questions are verbatim; answer options are abridged from the corpus answer key.

Source: ehs-ai-grounding-eval, corpus/items/pressure_relief.json and results/baselines_*.json. Scoring policy weights from the same run config. Try more items.

Nine of fifteen headline claims are verified, three are drafts, three are untested

Every claim carries the file that holds its evidence and how far it has been tested.

04

Fifteen claims, each with its evidence file and how far it was tested

Nine verified, three draft, three not yet tested. Filter by status or search the text.

Programme claims with evidence repository and status
ClaimEvidenceStatus
A plausibility screen raises the pooled aggregate TRIR from 0.134 to 3.983, a 29.7x ratio, over 2,801,064 filings.Regression tests assert the headline values against the data.ehs-osha-analysisVerified
2.07% of filings carry 96.68% of reported hours.Implausible means hours per employee outside 120 to 4,500; the bounds are a choice, with a sensitivity grid.ehs-osha-analysisVerified
The raw 2019 national TRIR is 249x below the screened one.Unscreened 0.016, screened 4.02, with 99.6% of hours in implausible filings.ehs-osha-analysisVerified
A flagged establishment is flagged again next year 27.2% of the time, against 0.95% for one never flagged (odds ratio 39.1).Measures what employers filed, not workplaces.ehs-osha-analysisVerified
Covariate-adjusted count models show NB2 beats Poisson in all 30 industries tested.ehs-osha-analysisVerified
Three non-LLM baselines on the 68-item benchmark: random floor 52.1%, TF-IDF 27.0%, oracle 100%.Baselines only. They say nothing about any language model.ehs-ai-grounding-evalVerified
The calculator's JavaScript and Python implementations produce identical numbers.Pinned by a test fixture. The ROI comparison is an accounting-scope argument, not a discovery.ehs-capitals-calculatorVerified
The ontology encodes IDHEAS-G's four contexts and twenty factors with a generated crosswalk and derivation traces.Taxonomy adopted, not new. Factor levels in worked scenarios are analyst-assigned.ehs-human-factors-ontologyVerified
197 citation keys cited in the bundled draft, 197 verified, 5 judged overstretched.groundedVerified
Grounded reasoning for safety-critical AI: ontology, benchmark, evaluation and data foundations in one manuscript.Not submitted. Four of five internal reviewers would reject it as it stands.groundedDraft
OSHA ITA data quality as a standalone paper: screen definition, sensitivity grid, and the nine-year series.Not submitted. The data-quality section is the only part of the programme with empirical findings.groundedDraft
SEM simulation studies show where the method breaks under rare events and misspecification.Every number comes from a known generating model. Hand-set path weights cannot be read causally.ehs-risk-semDraft
How often a language model substitutes an adjacent device when answering safety-critical questions.No language model has been run through the benchmark.ehs-ai-grounding-evalNot yet tested
Binding every claim to a source clause and anchor span prevents the substitution in a real system.The adapter and one-command run script exist; no result yet.ehs-ai-grounding-evalNot yet tested
SEM risk scoring on real injury data.The SEM work is entirely synthetic; no real injury data has been analysed with it.ehs-risk-semNot yet tested

Six repositories answer one failure each: one study, two apparatus, one simulation, two tools

They sort into three lanes: bind the answer, test the denominator, price the decision.

05

One empirical study, two apparatus, one simulation, two tools

Six repositories in three lanes. The chips say which is which, and each card links to its own five-minute story.

BIND THE ANSWER TEST THE DENOMINATOR PRICE THE DECISION grounding-eval 3 BASELINES ontology 20 FACTORS osha-analysis 2.8M FILINGS benchmarks MIT RATE LIBRARY risk-sem 0 REAL RECORDS capitals-calculator 2 IMPLEMENTATIONS ONE FAILURE CHECKED AGAINST BAD DATA

Six repositories, three lanes, one failure underneath them. The cards below carry the same six, each linking to its own story page.

The failure

An ungrounded answer, checked against untrustworthy data

  1. A raw OSHA benchmark can move 249x on a single plausibility screen, so any injury-rate benchmark you buy needs to state its screen.

    The chart →
  2. A site flagged for implausible hours one year is 39.1 times more likely to be flagged again the next (27.2% vs 0.95%).

    The persistence finding →
  3. The capitals calculator prices a safety investment both ways: booked cost only, and booked cost plus the human-capital cost that rarely gets charged.

    Open the calculator →

No language model has been run through the benchmark yet

Four things a reader should know before citing anything on this page.

06

Four things are not established, and they are listed here rather than buried

No model result, one rejected manuscript verdict, an adopted taxonomy, and a simulation with no real data behind it.

  • 01No language model has been run through the benchmark. Three non-LLM baselines exist (random floor 52.1%, TF-IDF 27.0%, oracle 100%); a real Anthropic Messages API adapter and a one-command run script now exist, but no model result should be inferred from the baselines.
  • 02Four of five internal reviewers would reject the bundled manuscript as it stands. The data-quality section is the only part with empirical findings.
  • 03The ontology's four contexts and twenty factors are IDHEAS-G's, taken unchanged and cited. Factor levels in worked scenarios are analyst-assigned.
  • 04The SEM work is entirely synthetic. Hand-set path weights cannot be read causally, and no real injury data has been analysed with it.

Dated changes across all six repositories: changelog (RSS feed).

Next · 1 of 7

2.8 million filings, and the screen that moves the national rate 249 times

The denominator, in full: what the hours screen flags, how much of the reported hours it removes, and why the same filers keep appearing. Five minutes.

Or press M for the site map.