GroundedThe argument
A wrong answer that survives review
Asked about a rupture disk, an assistant answered about a relief valve. Nothing signalled the swap. The public safety data meant to catch it is no more trustworthy.
How to read this site
One argument told in sequence, from the denominator to the decision. About 30 minutes.
Start the guided path → Analyst Project pages and toolsMethods, tables, code and the interactive tools behind every number on this page.
Open the six repositories → Reviewer Papers, claim ledger, what is missingFifteen headline claims with their evidence files, plus an honest list of what is not established.
Read the claim ledger →The pivot to data
One reporting year. One plausibility screen. No single multiplier repairs a benchmark built on raw OSHA filings.
Versus 0.95% for an establishment never flagged before.
Data quality, a grounding benchmark, an ontology, a simulation, a calculator, a paper trail. No language model has been scored on the benchmark yet.
Safety numbers are only as good as the denominator and the citation underneath them.Programme status: see changelog for current counts.
Five moves run from a wrong answer to one question you can ask on Monday
Each move states its finding and the number behind it. The figure on the left changes with the move.
Zero parameters are shared by the device asked about and the device answered about
A rupture disk question came back with set pressure, blowdown and reseating. Those are relief valve parameters; a rupture disk has none of them. Nothing in the answer contradicts itself, so a consistency check finds no tell.
0 shared parametersOne plausibility screen on hours moves the 2019 national rate 249 times
In reporting year 2019 the raw national TRIR is 0.016 and the screened one is 4.018, a ratio of 249.3. Across 2,801,064 real OSHA filings from 2016 to 2024 the raw series swings from 0.016 to 2.92 while the screened series stays between 3.15 and 4.60.
249x, 2019A flagged site is flagged again 27.2 percent of the time, against 0.95 percent for the rest
That is an odds ratio of 39.1. A permutation null centred at 1.0 rules out independent keying errors as the main mechanism. The bad hours sit with the same filers year after year, so the noise does not average out.
39.1odds ratioSix repositories answer one failure each, in three lanes
One is empirical, two are apparatus, one is simulation, two are tools. They bind the answer to a source, test the denominator under the rate, and price what the decision costs. Each one ships its own data and its own story page.
6repositoriesThree of the fifteen headline claims have not been tested at all
Nine of the fifteen are verified against files in the repositories and three are drafts. No language model has been run through the benchmark, the SEM work is entirely synthetic, and both papers are unsubmitted. All of that is in the ledger below, not in a footnote.
9 / 3 / 3verified / draft / untestedOne question separates a rate from a number: what is the hours screen?
On this data the screen removes 96.68 percent of reported hours. Ask any vendor or internal report for its hours-per-employee screen and for the share of hours that screen removes. If there is no screen, the rate is not a rate.
1questionA rupture disk question came back with relief valve parameters, and it read as competent
Two parameter spaces that do not intersect, and nothing in the wording signals the swap.
Asked about a rupture disk, answered about a relief valve
Fluent, well formatted, and inapplicable to the device in front of the engineer. The two devices share no parameters, so every value in the answer is plausible for the device it was not about.
Rupture disk
Opens once and stays open.
- Burst pressure
- Burst tolerance
- Manufacturing design range
- Operating ratio
Relief valve
Parameters of a device that recloses. A rupture disk has none of them.
- Set pressure
- Blowdown
- Reseating
- Chatter
Nothing in the answer is wrong on its own terms
The two parameter spaces do not intersect, so every value is a plausible value for a relief valve. A reviewer checking internal consistency finds none of the usual tells: no contradictions, no missing units, no hedging. The error lives in the binding between the question and the device, and that binding was never made.
Authoritative binding replaces free retrieval: every claim carries a source clause and an anchor span. A benchmark scores abstention as a correct response when the corpus does not contain the answer. The ontology makes the device and context explicit before any parameter is named. And the data work underneath asks whether the public benchmarks used to check any of this are themselves trustworthy.
- 1
No retrieval happened
The answer came from parametric memory, not a bound source.
- 2
Correct content is long tail
Relief valves dominate the training distribution. Disks do not.
- 3
Abstention is penalised
A confident answer scores better than "not in my sources".
- 4
The corpus is lexically adversarial
Standards share vocabulary across devices that share no parameters.
A plausibility screen on hours moves the 2019 national rate by a factor of 249
Any benchmark quoted off raw OSHA filings can be wrong by two orders of magnitude.
Screening hours turns a series that swings 0.016 to 2.92 into one that holds near four
The screen is a plausibility test on hours per employee. It flags 2.07% of filings, and those filings carry 96.68% of all reported hours.
Real public OSHA ITA data, 2016 to 2024.
carry 96.68% of all reported hours.
Unscreened to screened, all years pooled.
Median screened establishment TRIR is 2.47.
Aggregate TRIR by reporting year. Screening produces a stable series. The raw file does not.
A filing is implausible when reported hours per employee fall outside 120 to 4,500. Pooled across nine years that flags 2.07% of 2,801,064 filings, which carry 96.68% of reported hours. The pooled ratio of screened to unscreened TRIR is 29.7x; the median screened establishment TRIR is 2.47. The bounds are a defensible choice rather than a discovered truth, which is why a sensitivity grid ships with the repository. The analysis measures what employers filed, not workplaces.
The same filers keep getting flagged. Probability of a flag this year, by last year's status.
Odds ratio 39.15,776 of 21,243 establishments flagged again.
13,394 of 1,415,236 establishments newly flagged.
Each grid is 100 establishments. A site flagged for implausible hours one year is 39.1 times more likely to be flagged again the next. A permutation null centred at 1.0 (95% range 0.92 to 1.09) rules out independent year-to-year keying errors as the dominant mechanism.
The trap is a clause that reads correct for the wrong device
Three real items from the 68-item corpus. Answer them before you read how the baselines did.
TF-IDF retrieval scores 27.0 percent, below a 52.1 percent random floor
Pick an answer on three real items from the 68-item corpus, then see the co-hyponym trap and how the three non-LLM baselines scored the same item. Questions are verbatim; answer options are abridged from the corpus answer key.
Source: ehs-ai-grounding-eval, corpus/items/pressure_relief.json and results/baselines_*.json. Scoring policy weights from the same run config. Try more items.
Nine of fifteen headline claims are verified, three are drafts, three are untested
Every claim carries the file that holds its evidence and how far it has been tested.
Fifteen claims, each with its evidence file and how far it was tested
Nine verified, three draft, three not yet tested. Filter by status or search the text.
| Claim | Evidence | Status |
|---|---|---|
| A plausibility screen raises the pooled aggregate TRIR from 0.134 to 3.983, a 29.7x ratio, over 2,801,064 filings.Regression tests assert the headline values against the data. | ehs-osha-analysis | Verified |
| 2.07% of filings carry 96.68% of reported hours.Implausible means hours per employee outside 120 to 4,500; the bounds are a choice, with a sensitivity grid. | ehs-osha-analysis | Verified |
| The raw 2019 national TRIR is 249x below the screened one.Unscreened 0.016, screened 4.02, with 99.6% of hours in implausible filings. | ehs-osha-analysis | Verified |
| A flagged establishment is flagged again next year 27.2% of the time, against 0.95% for one never flagged (odds ratio 39.1).Measures what employers filed, not workplaces. | ehs-osha-analysis | Verified |
| Covariate-adjusted count models show NB2 beats Poisson in all 30 industries tested. | ehs-osha-analysis | Verified |
| Three non-LLM baselines on the 68-item benchmark: random floor 52.1%, TF-IDF 27.0%, oracle 100%.Baselines only. They say nothing about any language model. | ehs-ai-grounding-eval | Verified |
| The calculator's JavaScript and Python implementations produce identical numbers.Pinned by a test fixture. The ROI comparison is an accounting-scope argument, not a discovery. | ehs-capitals-calculator | Verified |
| The ontology encodes IDHEAS-G's four contexts and twenty factors with a generated crosswalk and derivation traces.Taxonomy adopted, not new. Factor levels in worked scenarios are analyst-assigned. | ehs-human-factors-ontology | Verified |
| 197 citation keys cited in the bundled draft, 197 verified, 5 judged overstretched. | grounded | Verified |
| Grounded reasoning for safety-critical AI: ontology, benchmark, evaluation and data foundations in one manuscript.Not submitted. Four of five internal reviewers would reject it as it stands. | grounded | Draft |
| OSHA ITA data quality as a standalone paper: screen definition, sensitivity grid, and the nine-year series.Not submitted. The data-quality section is the only part of the programme with empirical findings. | grounded | Draft |
| SEM simulation studies show where the method breaks under rare events and misspecification.Every number comes from a known generating model. Hand-set path weights cannot be read causally. | ehs-risk-sem | Draft |
| How often a language model substitutes an adjacent device when answering safety-critical questions.No language model has been run through the benchmark. | ehs-ai-grounding-eval | Not yet tested |
| Binding every claim to a source clause and anchor span prevents the substitution in a real system.The adapter and one-command run script exist; no result yet. | ehs-ai-grounding-eval | Not yet tested |
| SEM risk scoring on real injury data.The SEM work is entirely synthetic; no real injury data has been analysed with it. | ehs-risk-sem | Not yet tested |
No claims match that filter.
Six repositories answer one failure each: one study, two apparatus, one simulation, two tools
They sort into three lanes: bind the answer, test the denominator, price the decision.
One empirical study, two apparatus, one simulation, two tools
Six repositories in three lanes. The chips say which is which, and each card links to its own five-minute story.
Six repositories, three lanes, one failure underneath them. The cards below carry the same six, each linking to its own story page.
An ungrounded answer, checked against untrustworthy data
A raw OSHA benchmark can move 249x on a single plausibility screen, so any injury-rate benchmark you buy needs to state its screen.
The chart →A site flagged for implausible hours one year is 39.1 times more likely to be flagged again the next (27.2% vs 0.95%).
The persistence finding →The capitals calculator prices a safety investment both ways: booked cost only, and booked cost plus the human-capital cost that rarely gets charged.
Open the calculator →
A reproducible pipeline downloads 2.8M real OSHA filings, screens them, and is covered by 175 stdlib unittest tests.
The OSHA pipeline →Covariate-adjusted count models show NB2 beats Poisson in all 30 industries tested.
The count-model result →SEM simulation studies show where the method breaks, under rare events and misspecification, run with
The SEM studies →python3 simulations/run_all.py.The grounding benchmark scores a real system with
Run the harness →python3 -m grounding_eval.cli run --adapter anthropic --grounded.
Two papers: the bundled draft and the OSHA data-quality paper split out from it.
The bundled draft →197 citation keys cited, 197 verified, 5 judged overstretched.
The citation audit →Four of five internal reviewers would reject the combined manuscript, which is why the OSHA paper now stands alone.
The review outcome →Every section states what it does not yet establish.
What is missing →
Two write-ups
Grounded reasoning for safety-critical AI
197 verified referencesOntology, benchmark, evaluation and data foundations in one manuscript, with a list of what it does not yet establish.
Open the paper → StandaloneOSHA ITA data quality
249x in one yearThe empirical result on its own: screen definition, sensitivity grid, and the nine-year series behind the chart above.
Open the OSHA paper →No language model has been run through the benchmark yet
Four things a reader should know before citing anything on this page.
Four things are not established, and they are listed here rather than buried
No model result, one rejected manuscript verdict, an adopted taxonomy, and a simulation with no real data behind it.
- 01No language model has been run through the benchmark. Three non-LLM baselines exist (random floor 52.1%, TF-IDF 27.0%, oracle 100%); a real Anthropic Messages API adapter and a one-command run script now exist, but no model result should be inferred from the baselines.
- 02Four of five internal reviewers would reject the bundled manuscript as it stands. The data-quality section is the only part with empirical findings.
- 03The ontology's four contexts and twenty factors are IDHEAS-G's, taken unchanged and cited. Factor levels in worked scenarios are analyst-assigned.
- 04The SEM work is entirely synthetic. Hand-set path weights cannot be read causally, and no real injury data has been analysed with it.
Dated changes across all six repositories: changelog (RSS feed).
Next · 1 of 72.8 million filings, and the screen that moves the national rate 249 times
The denominator, in full: what the hours screen flags, how much of the reported hours it removes, and why the same filers keep appearing. Five minutes.
Or press M for the site map.