Try the denominator
Three ways into the same public filings. Move the plausibility window and watch what it does to the national rate, switch the yearly series, then place a site's TRIR against its industry peers. Full denominator explorer →
Move the screen
Hours per employee per year a filing must fall between to count.
TRIR by year
Hover or tab through the bars for the year.
Select a bar to read the year.
Source artifact: outputs/summary.json - quality.by_yearPeer percentile lookup
Pooled screened percentiles by three-digit NAICS and size band.
Why this exists
The correction is real, it is large, and it is not a constant. That last part is what stops it being a footnote.
Every establishment-level injury benchmark in the United States is computed from the same public filings analysed here: 2,801,064 of them, calendar years 2016 through 2024, downloaded from the OSHA Injury Tracking Application by a script in this repository. The rate everyone quotes has hours worked in its denominator.
Read on
Only 2.07% of filings report hours that cannot be right. Those filings carry 96.7% of every hour in the dataset. Pooled aggregate TRIR reads 0.134 before the screen and 3.983 after it.
A ratio of 29.7x would still be manageable if it held still. It does not. The same screen moves 2018 by 1.39x and 2019 by 249x. Nobody can apply a published correction factor to their own year and know what they have.
What follows from that
Read on
- An unscreened aggregate is not a noisy estimate of the screened one. Their relationship changes year to year, so the error has no fixed sign or size.
- The screened series is stable. That is the useful finding: screening does not just shift the number, it makes the series comparable across years at all.
- Anyone publishing a benchmark from these files owes readers the screen they used. Without it the number is not reproducible even in principle.
Findings
Four numbers, each traceable to the artifact named beneath it.
Aggregate TRIR computed with and without the plausibility screen differs by 1.39x in 2018 and 249x in 2019.Why
2.07% of filings fail the hours-per-employee screen, and they carry 96.7% of every hour reported.Why
Screened aggregate TRIR stays inside a narrow band across all 9 years, while the unscreened figure swings between 0.016 and 2.92.Why
The median screened establishment sits well below the screened aggregate, which is what a right-skewed count distribution looks like.Why
What this does and does not show
What this shows
- Which filings report hours that cannot be right: 2.07% of 2,801,064 filings, carrying 96.7% of all hours.
- That the screened-to-unscreened correction is not a constant. It runs 1.39x to 249x depending on the year.
- That the result does not hinge on where the window is drawn: across the sensitivity grid corrected TRIR stays between 3.963 and 3.993.
- That screened percentile ordering is stable between years, while the low bands move most.
What this does not show
- The screen identifies filings that cannot be right. It does not identify filings that are merely wrong, and it cannot recover the true hours for an excluded establishment.
- Aggregates after screening are conditional on the surviving population, which is not a random sample of the original. If implausible hours are filed disproportionately by one kind of employer, the screened aggregate inherits that selection.
- Nothing here is a claim about whether workplaces got safer. It is a claim about what the denominator will support.
Year by year
The pooled ratio is an average over years that do not resemble each other.Why
| Year | Filings | Flagged | Hours flagged | TRIR raw | TRIR screened | Ratio | Median est. TRIR |
|---|---|---|---|---|---|---|---|
| 2016 | 214,977 | 2.30% | 31.6% | 2.729 | 3.921 | 1.44x | 2.79 |
| 2017 | 259,757 | 1.66% | 35.2% | 2.634 | 3.998 | 1.52x | 2.94 |
| 2018 | 286,884 | 1.56% | 29.1% | 2.920 | 4.069 | 1.39x | 3.00 |
| 2019 | 290,475 | 1.60% | 99.6% | 0.016 | 4.018 | 249.28x | 2.93 |
| 2020 | 293,385 | 1.91% | 38.6% | 2.792 | 4.491 | 1.61x | 2.39 |
| 2021 | 315,936 | 2.60% | 46.1% | 2.385 | 4.368 | 1.83x | 2.53 |
| 2022 | 346,799 | 2.02% | 42.9% | 2.661 | 4.605 | 1.73x | 2.42 |
| 2023 | 394,231 | 2.34% | 56.7% | 1.380 | 3.147 | 2.28x | 2.00 |
| 2024 | 398,620 | 2.37% | 92.4% | 0.280 | 3.642 | 12.99x | 1.69 |
Why
Where the hours are
A national aggregate rate is a sum of cases over a sum of hours.Why
Hours concentration in the largest filings
Why
Figures
Two interactive charts from summary.json.Why
Screened TRIR stays between 3.1 and 4.6. Unscreened swings 0.02 to 2.92.
Source artifact: outputs/summary.json - quality.by_yearFlagged filings hold 29% to 100% of all hours, depending on the year.
Source artifact: outputs/summary.json - quality.by_yearStatic figures
Aggregate TRIR before and after the plausibility screen
Unscreened TRIR tracks bad hours data, not safety performance.
Why
A small share of filings carries most of the reported hours
Implausible filings hold from a third to nearly all reported hours, year to year.
Why
Hours worked per average employee (log10 scale)
The tail runs many orders of magnitude beyond anything a workforce can produce (screen: 120-4500 h).
Why
Zero-recordable share falls sharply with establishment size
A zero rate reflects headcount as much as safety performance.
Why
Peer-group 75th percentile TRIR: 2023 vs 2024
A benchmark only works if the peer band holds steady between years (469 matched peer groups).
Why
Count models
Rates are ratios of counts, and the count distribution decides what a control limit or a significance test on that rate is allowed to say.
| model | n industries best by aic | n industries fitted | share best by aic |
|---|---|---|---|
| poisson | 0 | 30 | 0.00% |
| nb2 | 15 | 30 | 50.00% |
| zip | 0 | 30 | 0.00% |
| zinb | 15 | 30 | 50.00% |
Notes
Poisson, negative binomial, zero-inflated Poisson and zero-inflated negative binomial were fitted per industry and compared by likelihood ratio and AIC. Boundary-corrected p-values are reported because the zero-inflation parameter sits on the edge of its parameter space, where the naive chi-squared reference distribution does not hold. Poisson never wins, which is the practical result: injury counts are overdispersed everywhere, so control limits built on a Poisson assumption are too tight.
Peer benchmarking
If a benchmark is going to be used to judge a site, the peer group it is judged against has to be stable between years.
| year from | year to | percentile | n matched groups | spearman rho | median abs change | median abs rel change |
|---|---|---|---|---|---|---|
| 2016 | 2017 | p10 | 406 | 0.8631 | 0 | 0.2704 |
| 2016 | 2017 | p25 | 406 | 0.9234 | 0.04302 | 0.1256 |
| 2016 | 2017 | p50 | 406 | 0.9096 | 0.1386 | 0.09389 |
| 2016 | 2017 | p75 | 406 | 0.8548 | 0.4369 | 0.09579 |
| 2016 | 2017 | p90 | 406 | 0.8867 | 0.9544 | 0.1016 |
| 2016 | 2017 | p95 | 406 | 0.9092 | 1.501 | 0.1271 |
| 2017 | 2018 | p10 | 444 | 0.8933 | 0 | 0.2584 |
| 2017 | 2018 | p25 | 444 | 0.9145 | 0.03925 | 0.1292 |
| 2017 | 2018 | p50 | 444 | 0.8913 | 0.1583 | 0.08632 |
| 2017 | 2018 | p75 | 444 | 0.8488 | 0.36 | 0.08182 |
| 2017 | 2018 | p90 | 444 | 0.8864 | 0.7765 | 0.08334 |
| 2017 | 2018 | p95 | 444 | 0.9143 | 1.1 | 0.09356 |
Why
| size band | n | zero share | mean recordables | variance to mean ratio | aggregate trir | median establishment trir |
|---|---|---|---|---|---|---|
| 001-019 | 535,187 | 75.11% | 0.388 | 3.351 | 4.939 | 0 |
| 020-049 | 884,997 | 46.22% | 1.252 | 6.459 | 4.582 | 2.546 |
| 050-099 | 553,610 | 23.92% | 3.054 | 7.851 | 4.977 | 3.35 |
| 100-249 | 493,919 | 10.90% | 6.143 | 11.38 | 4.707 | 3.608 |
| 250-499 | 170,749 | 7.30% | 12.3 | 14.74 | 4.132 | 3.382 |
| 500-999 | 61,329 | 8.19% | 20.42 | 29.54 | 3.375 | 2.381 |
| 1000+ | 43,416 | 8.13% | 72.25 | 241.4 | 3.143 | 2.844 |
Why
Method and limits
The screen, its sensitivity, and what the result does not license.
Method, limits and notes
The plausibility screen
Hours worked divided by average employees must fall between 120 and 4,500 per year. Below that, hours were filed in the wrong unit. Above it, a keying error. Filings failing the screen are excluded from aggregates rather than corrected, because the true value is not recoverable from the filing alone.
How sensitive is that window
The bounds are a judgement call, so the pipeline sweeps them. Across the grid the corrected aggregate TRIR moves only between 3.963 and 3.993, and the flagged share between 1.79% and 4.49%. The finding does not depend on where exactly the window is drawn.
Reproducing this
The pipeline downloads the public ITA files itself and fails loudly rather than substituting fixtures if they are unavailable. 4,703 duplicate rows were dropped before analysis. Every table and figure on this page is generated by script; none is hand-entered.
Notes
Reviewer note
An internal adversarial review held that this analysis is a separable contribution being buried inside a larger manuscript, and that it should stand as its own paper. That criticism is recorded here rather than answered.
Run it yourself
python scripts/download_data.py --list # show the catalog and URLs
python scripts/download_data.py # ~170 MB into data/raw
python scripts/run_analysis.py # writes outputs/
cd tests && python -m unittest discover -s . -p "test_*.py"Or make data, make analysis, make test. The pipeline downloads the public ITA files itself and fails loudly rather than substituting fixtures.