Start here
A safety number is only as good as the denominator under it, the source clause behind it, and the peer group it is ranked against.
This programme tested all three. Below are the six findings, in order, each with the figure it rests on and the file the figure was built from. Every stop has two doors: the five-minute story for a leader, and the project page for an analyst.
A single plausibility screen moves the 2019 national injury rate by 249 times.
One keying error in the hours column can take over the denominator of a national rate.
Aggregate TRIR across 2.8M OSHA 300A filings was computed twice: once on the filings as submitted, once after a screen on hours per employee. In 2019 the two answers are 0.016 and 4.018. The screened series stays between 3.15 and 4.61 in every year; the unfiltered series does not.
A site flagged for implausible hours one year is flagged again the next 27.2% of the time, against 0.95% for a site never flagged. The noise sits with the same filers.
On the grounding benchmark, retrieval scores below the random floor.
Picking the nearest text is worse than guessing when the nearest text is the adjacent clause.
The benchmark holds 68 items. Every item carries a source clause and an anchor span, and abstention counts as correct when the corpus does not hold the answer. Three baselines have been run: a random floor at 52.1%, TF-IDF retrieval at 27.0%, and an oracle ceiling at 100%.
Retrieval falls below the floor because it substitutes a neighbouring clause that reads correct. That is the failure the benchmark was built to catch.
Only 31 of 80 human-factors mappings are close matches.
Four established frameworks name the same 20 factors differently, and six pairs have no counterpart at all.
Twenty performance-influencing factors were encoded in Turtle and crosswalked to SPAR-H, CREAM, HFACS and the HSE performance-influencing factors. Of the 80 factor-to-framework pairs, 31 are close, 23 are broader, 20 are partial and 6 have no counterpart.
HFACS carries no procedures category at the precondition level, so that factor maps upward to an organisational process instead. Each mapping records its strength and its reference, so a reviewer can see the judgment rather than infer it.
The same path model needs 153 cases to pass a fit test and 6,884 to separate two coefficients.
A model can clear every fit index it is asked about and still be far too small to answer the question it was built for.
Simulation studies sized the same safety-risk path model against five different claims. A test of close fit reaches 80% power at 153 cases. A single standardised coefficient reaches a standard error of 0.05 at 369 cases, and 0.025 at 1,459.
Telling two coefficients apart is the expensive claim. Separating 0.45 from 0.30 takes 770 cases; separating -0.25 from -0.20 takes 6,884. Sample sizes in this range are the ones practitioners are usually quoting on.
Pricing the uncounted costs cuts the payback period from 3.41 years to 1.59.
The control did not change. The accounting did.
One $150,000 loading-station control was priced twice. The booked ledger counts downtime, spilled material and injury cost: three invoiced lines totalling $55,000 a year. It pays back in 3.41 years.
Lost work time, retraining after turnover, productivity loss and the effect on the crew add $63,000 a year that no invoice carries. On that ledger the same control pays back in 1.59 years. Which ledger a committee uses decides the answer.
The same 3.20 injury rate ranks at the 76th percentile of its industry and the 93rd among its own size band.
The peer group, not the rate, decides whether a site looks safe.
A site with 600,000 hours and 10 recordable cases reports a TRIR of 3.20. Against every establishment that filed a 300A, that is 8% below the all-filer aggregate of 3.49, which is the version that reaches the board deck.
Hold the industry and it ranks p76 among the 478 filers in NAICS 3252. Add the size band and it ranks p93 among 34 peers. Same rate, three verdicts.
Any group that small gets its n printed next to the percentile, not a headline.
Pick your path
Three ways through the same material. Take the one that matches what you have to decide.
Read the six stories in order
One argument across six sites, each a five-minute read with one figure and one action. Start with the denominator and follow the bar at the bottom of every page.
Start the first story →Open the project pages and tools
Tables, pipelines and the interactive tools: the filing explorer, the benchmark item runner, the ontology walkthrough and crosswalk, the cost calculator and the MIT-licensed rate library.
Open the Observatory →Read the drafts and the claim ledger
Two unsubmitted drafts, a ledger that marks each claim tested, partly tested or not yet tested, and a plain list of what this programme has not established.
Open the bundled draft →