Six minutes · six stops

Start here

A safety number is only as good as the denominator under it, the source clause behind it, and the peer group it is ranked against.

This programme tested all three. Below are the six findings, in order, each with the figure it rests on and the file the figure was built from. Every stop has two doors: the five-minute story for a leader, and the project page for an analyst.

Searching

A single plausibility screen moves the 2019 national injury rate by 249 times.

One keying error in the hours column can take over the denominator of a national rate.

249x2019 screened rate against the rate as filed

Aggregate TRIR across 2.8M OSHA 300A filings was computed twice: once on the filings as submitted, once after a screen on hours per employee. In 2019 the two answers are 0.016 and 4.018. The screened series stays between 3.15 and 4.61 in every year; the unfiltered series does not.

A site flagged for implausible hours one year is flagged again the next 27.2% of the time, against 0.95% for a site never flagged. The noise sits with the same filers.

Aggregate TRIR before and after the plausibility screen, 2016 to 2024 Screened TRIR stays between 3.15 and 4.61 across nine years. Unscreened TRIR collapses to 0.016 in 2019 and 0.280 in 2024. 0 1 2 3 4 5 16 17 18 19 20 21 22 23 24 2019: 249x apart Screened Unscreened, as filed
Source: ehs-osha-analysis, outputs/tables/quality_by_year.csv. Persistence figures from outputs/tables/panel_transitions.csv and docs/PANEL.md.

On the grounding benchmark, retrieval scores below the random floor.

Picking the nearest text is worse than guessing when the nearest text is the adjacent clause.

27.0%TF-IDF retrieval, against a 52.1% random floor

The benchmark holds 68 items. Every item carries a source clause and an anchor span, and abstention counts as correct when the corpus does not hold the answer. Three baselines have been run: a random floor at 52.1%, TF-IDF retrieval at 27.0%, and an oracle ceiling at 100%.

Retrieval falls below the floor because it substitutes a neighbouring clause that reads correct. That is the failure the benchmark was built to catch.

No language model evaluatedBaselines only. No model result should be read out of them.
Three non-language-model baselines on the 68 item grounding benchmark Random floor scores 52.1 percent, TF-IDF retrieval scores 27.0 percent, the oracle ceiling scores 100 percent. No language model has been scored. Random floor 52.1% TF-IDF retrieval 27.0% Oracle ceiling 100.0% Retrieval sits below the floor 0100% of 68 items
Source: ehs-ai-grounding-eval, results/baselines_*.json and corpus/items/. Scoring policy weights from the same run config.

Only 31 of 80 human-factors mappings are close matches.

Four established frameworks name the same 20 factors differently, and six pairs have no counterpart at all.

31 / 80close matches across four frameworks

Twenty performance-influencing factors were encoded in Turtle and crosswalked to SPAR-H, CREAM, HFACS and the HSE performance-influencing factors. Of the 80 factor-to-framework pairs, 31 are close, 23 are broader, 20 are partial and 6 have no counterpart.

HFACS carries no procedures category at the precondition level, so that factor maps upward to an organisational process instead. Each mapping records its strength and its reference, so a reviewer can see the judgment rather than infer it.

Match strength of 20 factors against four human factors frameworks Across 80 factor to framework pairs, 31 are close matches, 23 broader, 20 partial and 6 have no counterpart. 6 2 7 5 SPAR-H 6 8 6 CREAM 6 8 5 HFACS 13 5 2 HSE PIFs Close Broader Partial No counterpart 20 factors per framework, 80 pairs in all
Source: ehs-human-factors-ontology, docs/crosswalk-data.json (80 cells, 96 recorded matches). Taxonomy adopted from IDHEAS-G, not new.

The same path model needs 153 cases to pass a fit test and 6,884 to separate two coefficients.

A model can clear every fit index it is asked about and still be far too small to answer the question it was built for.

6,884cases to tell -0.25 from -0.20 at 80% power

Simulation studies sized the same safety-risk path model against five different claims. A test of close fit reaches 80% power at 153 cases. A single standardised coefficient reaches a standard error of 0.05 at 369 cases, and 0.025 at 1,459.

Telling two coefficients apart is the expensive claim. Separating 0.45 from 0.30 takes 770 cases; separating -0.25 from -0.20 takes 6,884. Sample sizes in this range are the ones practitioners are usually quoting on.

Synthetic dataEvery number here comes from a known generating model. No real records were used.
Sample size required for five claims about the same path model A global fit test needs 153 cases. A single coefficient precise to 0.05 needs 369. Separating a coefficient of minus 0.25 from minus 0.20 needs 6884. 100 1,000 10,000 80% power for close fit 153 SE = 0.05 on one path 369 Separate 0.45 from 0.30 770 SE = 0.025 on one path 1,459 Separate -0.25 from -0.20 6,884 Cases required, log scale
Source: ehs-risk-sem, results/study01_requirements.csv. Bases listed in the same file: RMSEA power at df = 80, standardised regression SE at R squared 0.288 and VIF 1.28.

Pricing the uncounted costs cuts the payback period from 3.41 years to 1.59.

The control did not change. The accounting did.

1.59 yrpayback once human-capital lines are priced

One $150,000 loading-station control was priced twice. The booked ledger counts downtime, spilled material and injury cost: three invoiced lines totalling $55,000 a year. It pays back in 3.41 years.

Lost work time, retraining after turnover, productivity loss and the effect on the crew add $63,000 a year that no invoice carries. On that ledger the same control pays back in 1.59 years. Which ledger a committee uses decides the answer.

Payback period for one loading station control under two accountings The same 150,000 dollar control pays back in 3.41 years on the booked ledger and 1.59 years once human capital lines are priced. 0 1 2 3 4 Booked ledger only 3.41 yr Human-capital lines priced 1.59 yr Same control. Same station. Two accountings. Years to payback
Source: ehs-capitals-calculator, the worked loading-station example in the calculator. JavaScript and Python implementations pinned to identical numbers by a test fixture.

The same 3.20 injury rate ranks at the 76th percentile of its industry and the 93rd among its own size band.

The peer group, not the rate, decides whether a site looks safe.

p93rank of a 3.20 TRIR among 250 to 499 employee sites in NAICS 3252

A site with 600,000 hours and 10 recordable cases reports a TRIR of 3.20. Against every establishment that filed a 300A, that is 8% below the all-filer aggregate of 3.49, which is the version that reaches the board deck.

Hold the industry and it ranks p76 among the 478 filers in NAICS 3252. Add the size band and it ranks p93 among 34 peers. Same rate, three verdicts.

Any group that small gets its n printed next to the percentile, not a headline.

Percentile rank of one TRIR of 3.20 against two peer groups A TRIR of 3.20 ranks at the 76th percentile among the 478 establishments in NAICS 3252 and at the 93rd percentile among the 34 establishments in that industry with 250 to 499 employees. One rate of 3.20. Two peer groups. NAICS 3252, all sizes (n=478) p76 NAICS 3252, 250 to 499 employees (n=34) p93 p0 p100 of the peer group 8% below the 3.49 all-filer aggregate.
Source: ehs-benchmarks, docs/data/benchmarks/32.json (keys 3252 and 3252|250-499) and the all-filer aggregate of 3.49 over 375,635 CY2025 filings. The 3.20 site is an illustrative input; every group it is compared against is real.
Where to go next

Pick your path

Three ways through the same material. Take the one that matches what you have to decide.

Leader · 30 minutes

Read the six stories in order

One argument across six sites, each a five-minute read with one figure and one action. Start with the denominator and follow the bar at the bottom of every page.

Start the first story →
Analyst · the methods

Open the project pages and tools

Tables, pipelines and the interactive tools: the filing explorer, the benchmark item runner, the ontology walkthrough and crosswalk, the cost calculator and the MIT-licensed rate library.

Open the Observatory →
Reviewer · the limits

Read the drafts and the claim ledger

Two unsubmitted drafts, a ledger that marks each claim tested, partly tested or not yet tested, and a plain list of what this programme has not established.

Open the bundled draft →