Empirical analysis - real public data

What this meansof every reported hour sits in the 2.07% of filings that fail the plausibility screen. Correcting for them moves aggregate TRIR by 1.39x in one year and 249x in another.

Source artifact: outputs/summary.json - quality.pooled

The hours column decides every benchmark built on it

A reproducible pipeline over 2,801,064 filings, CY2016-2024. 2.07% of filings carry 96.7% of all hours, and the correction ranges 1.39x to 249x by year. It cannot be published as a constant.

What this does not establishNothing here says a flagged filing is wrong, or that any establishment is unsafe. The screen marks hours that cannot be reconciled with the headcount reported beside them. It does not identify a cause, and it does not correct the record.

Filings2,801,064
Years2016-2024
Flagged2.07%
Of all hours96.7%

Try the denominator

Three ways into the same public filings. Move the plausibility window and watch what it does to the national rate, switch the yearly series, then place a site's TRIR against its industry peers. Full denominator explorer →

A

Move the screen

Hours per employee per year a filing must fall between to count.

Filings flagged-
Hours in flagged filings-
TRIR screened-
TRIR unscreened-
Ratio-

Source artifact: outputs/tables/quality_sensitivity_grid.csv (pooled, all years)
B

TRIR by year

Hover or tab through the bars for the year.

Select a bar to read the year.

Source artifact: outputs/summary.json - quality.by_year
C

Peer percentile lookup

Pooled screened percentiles by three-digit NAICS and size band.

Source artifact: outputs/tables/peer_percentiles_pooled.csv

Why this exists

The correction is real, it is large, and it is not a constant. That last part is what stops it being a footnote.

Every establishment-level injury benchmark in the United States is computed from the same public filings analysed here: 2,801,064 of them, calendar years 2016 through 2024, downloaded from the OSHA Injury Tracking Application by a script in this repository. The rate everyone quotes has hours worked in its denominator.

Read on

Only 2.07% of filings report hours that cannot be right. Those filings carry 96.7% of every hour in the dataset. Pooled aggregate TRIR reads 0.134 before the screen and 3.983 after it.

A ratio of 29.7x would still be manageable if it held still. It does not. The same screen moves 2018 by 1.39x and 2019 by 249x. Nobody can apply a published correction factor to their own year and know what they have.

What follows from that

Read on
  • An unscreened aggregate is not a noisy estimate of the screened one. Their relationship changes year to year, so the error has no fixed sign or size.
  • The screened series is stable. That is the useful finding: screening does not just shift the number, it makes the series comparable across years at all.
  • Anyone publishing a benchmark from these files owes readers the screen they used. Without it the number is not reproducible even in principle.

Findings

Four numbers, each traceable to the artifact named beneath it.

249×
Worst-year divergence

Aggregate TRIR computed with and without the plausibility screen differs by 1.39x in 2018 and 249x in 2019.

Why
A correction that varies by two orders of magnitude between adjacent years is not a portable constant.

summary.json - quality.by_year
96.7%
Hours concentrated in flagged filings

2.07% of filings fail the hours-per-employee screen, and they carry 96.7% of every hour reported.

Why
Aggregate TRIR reads 0.134 unscreened against 3.983 screened.

summary.json - quality.pooled
3.1-4.6
Screening yields a stable series

Screened aggregate TRIR stays inside a narrow band across all 9 years, while the unscreened figure swings between 0.016 and 2.92.

Why
The screen is doing more than trimming outliers.

outputs/tables/quality_by_year.csv
2.47
Median establishment TRIR

The median screened establishment sits well below the screened aggregate, which is what a right-skewed count distribution looks like.

Why
Reporting a mean against this shape overstates the typical site.

summary.json - quality.pooled

What this does and does not show

What this shows

  • Which filings report hours that cannot be right: 2.07% of 2,801,064 filings, carrying 96.7% of all hours.
  • That the screened-to-unscreened correction is not a constant. It runs 1.39x to 249x depending on the year.
  • That the result does not hinge on where the window is drawn: across the sensitivity grid corrected TRIR stays between 3.963 and 3.993.
  • That screened percentile ordering is stable between years, while the low bands move most.

What this does not show

  • The screen identifies filings that cannot be right. It does not identify filings that are merely wrong, and it cannot recover the true hours for an excluded establishment.
  • Aggregates after screening are conditional on the surviving population, which is not a random sample of the original. If implausible hours are filed disproportionately by one kind of employer, the screened aggregate inherits that selection.
  • Nothing here is a claim about whether workplaces got safer. It is a claim about what the denominator will support.

Year by year

The pooled ratio is an average over years that do not resemble each other.

Why
The point of this table is the spread, not the centre.

YearFilingsFlaggedHours flaggedTRIR rawTRIR screenedRatioMedian est. TRIR
2016214,9772.30%31.6%2.7293.9211.44x2.79
2017259,7571.66%35.2%2.6343.9981.52x2.94
2018286,8841.56%29.1%2.9204.0691.39x3.00
2019290,4751.60%99.6%0.0164.018249.28x2.93
2020293,3851.91%38.6%2.7924.4911.61x2.39
2021315,9362.60%46.1%2.3854.3681.83x2.53
2022346,7992.02%42.9%2.6614.6051.73x2.42
2023394,2312.34%56.7%1.3803.1472.28x2.00
2024398,6202.37%92.4%0.2803.64212.99x1.69
Read across the last two columns rather than down them.
Why
The screened aggregate and the median establishment barely move; the ratio between screened and unscreened moves by a factor of nearly two hundred.
Source artifact: outputs/summary.json - quality.by_year

Where the hours are

A national aggregate rate is a sum of cases over a sum of hours.

Why
When one filing dominates the second sum, it decides the answer on its own.

Hours concentration in the largest filings

Top 1 filing88.81%Top 10 filings94.54%Top 100 filings95.39%Top 1,000 filings96.23%Top 10,000 filings97.04%
One keying error in the hours column sets the national denominator.
Why
A single establishment filing accounts for 88.8% of the 19.0 trillion hours in the pooled dataset while contributing 0.00% of the recordable cases. Share of all reported hours held by the largest filings, ranked by hours.
Source artifact: outputs/tables/quality_hours_concentration.csv

Figures

Two interactive charts from summary.json.

Why
The pipeline's static figures sit underneath.

Screened TRIR stays between 3.1 and 4.6. Unscreened swings 0.02 to 2.92.

Source artifact: outputs/summary.json - quality.by_year

Flagged filings hold 29% to 100% of all hours, depending on the year.

Source artifact: outputs/summary.json - quality.by_year
Static figures
fig01_aggregate_trir_by_year.svg

Aggregate TRIR before and after the plausibility screen

Unscreened TRIR tracks bad hours data, not safety performance.

0 1 2 3 4 5 unscreened (all filings) screened 2016 2017 2018 2019 2020 2021 2022 2023 2024 2019 unscreened 0.02 2019 screened 4.02 2024 unscreened 0.28 2024 screened 3.64 Reporting year Recordable cases per 200,000 hours
The unscreened series is governed by how much bad hours data entered that year's file, not by safety performance.
Why
That is why the two series cross and diverge rather than tracking each other: the unscreened aggregate is not a noisy version of the screened one. Aggregate TRIR by year, screened against unscreened.
Source artifact: outputs/figures/fig01_aggregate_trir_by_year.svg
fig02_hours_share_implausible.svg

A small share of filings carries most of the reported hours

Implausible filings hold from a third to nearly all reported hours, year to year.

0 20 40 60 80 100 % of filings flagged implausible % of all reported hours held by them 2016 2017 2018 2019 2020 2021 2022 2023 2024 Reporting year Percent
Implausible filings hold anywhere from roughly a third to nearly all of the reported hours, depending on the year.
Why
That swing is what drives the instability in the ratio above. Share of all reported hours sitting in filings that fail the plausibility screen, by year.
Source artifact: outputs/figures/fig02_hours_share_implausible.svg
fig03_hours_per_employee.svg

Hours worked per average employee (log10 scale)

The tail runs many orders of magnitude beyond anything a workforce can produce (screen: 120-4500 h).

1 10 100 1000 10000 100000 1e+06 0 1 2 3 4 5 6 7 8 log10(annual hours per average employee) Number of filings (log scale)
The tail extends many orders of magnitude beyond anything a workforce can produce, so values outside the window are unit or keying errors rather than real labour.
Why
Distribution of hours worked per employee; the screen retains 120 to 4,500 hours per employee per year.
Source artifact: outputs/figures/fig03_hours_per_employee.svg
fig05_zero_share_by_size.svg

Zero-recordable share falls sharply with establishment size

A zero rate reflects headcount as much as safety performance.

0 0.2 0.4 0.6 0.8 1 001-019 020-049 050-099 100-249 250-499 500-999 1000+ Establishment size band (annual average employees) Share of establishments reporting zero recordables
A zero rate says as much about headcount as about safety performance: zero is the modal outcome at small establishments.
Why
Share of establishments reporting zero recordable cases, by size band.
Source artifact: outputs/figures/fig05_zero_share_by_size.svg
fig06_percentile_stability.svg

Peer-group 75th percentile TRIR: 2023 vs 2024

A benchmark only works if the peer band holds steady between years (469 matched peer groups).

0 5 10 15 20 25 0 5 10 15 20 25 p75 TRIR in 2023 p75 TRIR in 2024
A benchmark is only usable if the band a site is compared against holds still enough between years to mean the same thing.
Why
Year-over-year stability of peer-group percentile bands.
Source artifact: outputs/figures/fig06_percentile_stability.svg

Count models

Rates are ratios of counts, and the count distribution decides what a control limit or a significance test on that rate is allowed to say.

modeln industries best by aicn industries fittedshare best by aic
poisson0300.00%
nb2153050.00%
zip0300.00%
zinb153050.00%
Which count model wins on AIC, across the thirty largest three-digit NAICS industries fitted for the model year.Source artifact: outputs/tables/count_model_selection_summary.csv
Notes

Poisson, negative binomial, zero-inflated Poisson and zero-inflated negative binomial were fitted per industry and compared by likelihood ratio and AIC. Boundary-corrected p-values are reported because the zero-inflation parameter sits on the edge of its parameter space, where the naive chi-squared reference distribution does not hold. Poisson never wins, which is the practical result: injury counts are overdispersed everywhere, so control limits built on a Poisson assumption are too tight.

Peer benchmarking

If a benchmark is going to be used to judge a site, the peer group it is judged against has to be stable between years.

year fromyear topercentilen matched groupsspearman rhomedian abs changemedian abs rel change
20162017p104060.863100.2704
20162017p254060.92340.043020.1256
20162017p504060.90960.13860.09389
20162017p754060.85480.43690.09579
20162017p904060.88670.95440.1016
20162017p954060.90921.5010.1271
20172018p104440.893300.2584
20172018p254440.91450.039250.1292
20172018p504440.89130.15830.08632
20172018p754440.84880.360.08182
20172018p904440.88640.77650.08334
20172018p954440.91431.10.09356
Rank correlation and median absolute change for each percentile band between consecutive years, over peer groups matched in both years.
Why
Ordering is stable: the median rank correlation runs from 0.86 to 0.93 across bands. The band values themselves are less so, and least so at the bottom, where p10 moves a median of 26% year over year against 10% for p75. Low percentiles sit near zero, so a small absolute move is a large relative one, which is exactly the region where a site is judged to have improved. Showing 12 of 48 rows.
Source artifact: outputs/tables/percentile_band_stability.csv
size bandnzero sharemean recordablesvariance to mean ratioaggregate trirmedian establishment trir
001-019535,18775.11%0.3883.3514.9390
020-049884,99746.22%1.2526.4594.5822.546
050-099553,61023.92%3.0547.8514.9773.35
100-249493,91910.90%6.14311.384.7073.608
250-499170,7497.30%12.314.744.1323.382
500-99961,3298.19%20.4229.543.3752.381
1000+43,4168.13%72.25241.43.1432.844
Establishment size against zero share, dispersion and rate.
Why
The variance-to-mean ratio rises steadily with size, so a single Poisson assumption cannot serve every band.
Source artifact: outputs/tables/size_band_effects.csv

Method and limits

The screen, its sensitivity, and what the result does not license.

Method, limits and notes

The plausibility screen

Hours worked divided by average employees must fall between 120 and 4,500 per year. Below that, hours were filed in the wrong unit. Above it, a keying error. Filings failing the screen are excluded from aggregates rather than corrected, because the true value is not recoverable from the filing alone.

How sensitive is that window

The bounds are a judgement call, so the pipeline sweeps them. Across the grid the corrected aggregate TRIR moves only between 3.963 and 3.993, and the flagged share between 1.79% and 4.49%. The finding does not depend on where exactly the window is drawn.

Reproducing this

The pipeline downloads the public ITA files itself and fails loudly rather than substituting fixtures if they are unavailable. 4,703 duplicate rows were dropped before analysis. Every table and figure on this page is generated by script; none is hand-entered.

Notes

Reviewer note

An internal adversarial review held that this analysis is a separable contribution being buried inside a larger manuscript, and that it should stand as its own paper. That criticism is recorded here rather than answered.

Run it yourself

ehs-osha-analysis
python scripts/download_data.py --list     # show the catalog and URLs
python scripts/download_data.py            # ~170 MB into data/raw
python scripts/run_analysis.py             # writes outputs/
cd tests && python -m unittest discover -s . -p "test_*.py"

Or make data, make analysis, make test. The pipeline downloads the public ITA files itself and fails loudly rather than substituting fixtures.