Story 1 of 7 · five scenes

The denominator

A safety rate is a fraction. Almost every argument about it happens in the numerator. In 2.8 million OSHA filings, the bottom half of the fraction is where the damage is.

Every number here is read from a committed file in this repository, named under the figure it sits in. Press P for presenter mode, M for the site map.

2% of filings carry 97% of the hours

57,857 filings in the public OSHA panel cannot be true, and almost every reported hour sits inside them. A rate built on the raw panel is a rate about those rows.

2,801,064 establishment-year filings
2016 to 2024, after 4,703 duplicate rows dropped
n = 2,801,064 filings. Source: outputs/tables/quality_overall.csv, quality_flag_prevalence.csv
Scene 01 / step 1

2,801,064 filings, nine years, not a sample

2,801,064 establishment-year filings from 2016 through 2024, after 4,703 duplicate rows were dropped. Each dot stands for about 2,800 of them. This is the whole public Injury Tracking Application panel.

Scene 01 / step 2

57,857 of them cannot be true

57,857 filings, 2.07% of the panel, fail an arithmetic screen: hours worked divided by average employees has to land between 120 and 4,500 a year, hours and employees have to be present, counts cannot be negative. By count that is a rounding error.

Scene 01 / step 3

Those 57,857 filings hold 96.68% of the hours

One failure mode does it: 24,617 filings report more than 4,500 hours per employee and carry 96.65% of all hours. Someone put a company's payroll hours in the box meant for one site. Quote a rate off the raw panel and you are quoting those rows.

In 2019 the screen moves TRIR by a factor of 249

The same filings give 0.016 unscreened and 4.02 screened. The low answer is the one that looks like a record year on a slide.

3.98 pooled screened TRIR
cases per 200,000 hours, 2016 to 2024
n = 9 years, all filers. Source: outputs/tables/quality_by_year.csv, metrics_by_year.csv, quality_sensitivity.csv
Scene 02 / step 1

Screened, TRIR runs 3.15 to 4.60

Aggregate TRIR on filings that pass the screen stays between 3.15 and 4.60 across the nine years, pooled at 3.98. That is the kind of series you would put on a slide.

Scene 02 / step 2

Unscreened, the pooled rate is 0.13

The same filings with no screen give a pooled TRIR of 0.13, a thirtyfold difference. The shaded gap is not an adjustment or a modelling choice. It is hours that do not exist inflating a denominator.

Scene 02 / step 3

2019 reads 0.016 instead of 4.02

In 2019, 99.60% of all reported hours sat in flagged filings, so unscreened TRIR reads 0.016 against 4.02 screened. The correction factor is 249x, and 2024 is next worst at 13.0x. Neither year announces itself; both look like record years.

Scene 02 / step 4

Move the bounds 25 ways and the answer moves 0.03

The 120 to 4,500 bounds are a judgement call, so the pipeline sweeps 25 variants. Corrected pooled TRIR moves only between 3.963 and 3.993 while the flagged share moves between 1.79% and 4.49%. Where you draw the line changes who gets excluded, not the answer.

A flagged site is flagged again 27% of the time

Clean sites fail at 0.95%. Bad hours belong to the filer, so an entry rule at a few hundred sites beats an annual cleanup.

1,436,479 establishment-year pairs
each restricted to sites that filed in both years
n = 1,436,479 pairs. Source: outputs/tables/panel_transitions.csv (pooled row), docs/PANEL.md
Scene 03 / step 1

1,436,479 site-year pairs separate typos from filers

If implausible hours are independent keying errors, a flag this year says nothing about next year. If certain establishments misreport systematically, the flags cluster. 1,436,479 establishment-year pairs, each restricted to sites that filed in both years, tell the two apart.

Scene 03 / step 2

Flagged once, flagged again 27.2% of the time

5,776 of 21,243 flagged establishment-years were flagged again the following year. Hold that number against the base rate in the next step.

Scene 03 / step 3

Clean sites fail at 0.95%, an odds ratio of 39.1

13,394 of 1,415,236 clean establishment-years were flagged the next year. That is an odds ratio of 39.1 with Cohen's kappa 0.276, and all eight year pairs land between 24.7 and 63.1. Kappa 0.28 is the honest half: most flags in a given year are still one-offs, so fixing the repeat filers fixes the hours, not every flag.

Chance gives an odds ratio of 1.01, the panel gives 39.1

200 random reshuffles of the flags never came within a factor of 35 of the observed value. The clustering is real.

200 permutations, seed 20260909
flag counts held fixed within each year
n = 200 permutations. Source: outputs/tables/panel_permutation.csv (seed 20260909)
Scene 04 / step 1

200 reshuffles test the typo explanation

Hold each year's count of flagged filings fixed, scatter those flags at random among that year's filers, and recompute the pooled persistence odds ratio. 200 draws, seed 20260909. That is the typo hypothesis made explicit.

Scene 04 / step 2

Chance answers 1.012

The permutation mean is 1.012 with a 95% interval of 0.924 to 1.092, centred where it should be if flags were independent year to year. The 200 individual draws are not stored in the repository, so this figure shows the interval the file reports rather than a fabricated histogram.

Scene 04 / step 3

The panel answers 39.1, and 0 of 200 reshuffles reached it

39.1 against a null of 1.012 is a different order of magnitude, not a borderline result that needs a p-value convention to survive. Bad hours are a property of certain filers, not of keyboards. That makes them fixable at the point of entry.

Poisson wins 0 of 30 industries

Control limits on injury counts assume Poisson. On this panel the assumption fails in every large industry tested, so the limits come out too tight.

0 of 30 industries won by Poisson
30 largest three-digit NAICS groups, 2024 recordable counts
n = 30 industries. Source: outputs/tables/count_model_selection_summary.csv, count_model_covariates_summary.csv
Scene 05 / step 1

Poisson wins 0 of 30 industries

2024 recordable counts in the 30 largest three-digit NAICS groups with at least 500 screened establishments, fitted with Poisson, negative binomial and the zero-inflated version of each, compared on AIC. Intercept-only, NB2 and ZINB split the 30 evenly at 15 each. Poisson takes none.

Scene 05 / step 2

With structure in the mean, NB2 wins 20 of 30

Add size-band and four-digit sub-industry fixed effects with a log-hours offset and NB2 wins 20, ZINB 10: part of the apparent excess zeros was mean heterogeneity. Overdispersion does not move. NB2 beats Poisson in all 30 industries at a boundary-corrected p below 0.001, median dispersion 0.785.

Scene 05 / step 3

Poisson control limits are too tight in all 30

Control charts, run charts and significance tests on injury counts are usually built on a Poisson assumption that holds in none of these 30 industries. Median dispersion is 0.785, so the limits come out too tight and sites get flagged for a quarter that was ordinary variance. Check the distribution before you act on a signal.

06

What this does not establish

The screen excludes, it does not repair

Failing filings are dropped from aggregates because the true hours cannot be recovered from the filing alone. The persistence result suggests per-establishment repair might be possible; no repair has been attempted or validated.

The establishment key is a location, not a company

313,527 of 1,238,236 establishment IDs (25.3%) appear with more than one EIN across years, and 108,855 (8.8%) with more than one company name. State is stable for 99.7% of IDs. The panel result reads as a property of physical sites and their filing practice, not of corporate entities.

Nothing here is causal

Size-band and sub-industry coefficients are descriptive contrasts inside a screened administrative panel. Excess zeros cannot be split between incident-free operation and under-recording; the model measures the excess, not its source.

The BLS comparison was not made

The intended chart against the BLS SOII published rate is missing. The BLS series catalog returned HTTP 403 from this machine, so no series ID could be confirmed, and no published rate was typed in from memory. The attempt is logged in data/bls/PROVENANCE.txt.

Coverage is filers, not industry

The ITA panel covers establishments required to submit electronically: larger, and in higher-hazard industries. The 30 fitted NAICS groups were chosen on size alone and are a statement about large industries, not about every industry.

The day-count fields are unscreened

The screen acts on hours and employees only. Severity rates stay exposed to mistyped day totals; the swings in screened severity in 2022 and 2024 are a reason to build a day-count screen, not a finding about injuries.

07

Monday

Screen the hours before you quote the rate.

Take the establishment-year rows behind your last TRIR number. Divide hours worked by average employees. Anything outside 120 to 4,500 is not a data-quality footnote, it is a filing you cannot compute a rate from. Count how many there are, then count what share of your total hours they carry. If those two numbers disagree the way they do here, the rate you presented was about the bad rows.

Then check whether the same sites fail twice. If they do, the fix is a validation rule at entry and a conversation with whoever owns that site's payroll export, not an annual data cleanup.

Do it against this repository
python scripts/download_data.py
python scripts/run_analysis.py
python3 -m ehs_osha.panel

Writes outputs/tables/quality_*.csv and panel_*.csv. The pipeline downloads the public files itself and fails loudly rather than substituting fixtures.