On this page
Draft manuscript - not submitted

The Hours Denominator: Data Quality in OSHA Establishment Injury Filings and What It Does to Every Benchmark Built on Them

A working draft. Four artifacts: a structural account of semantically adjacent substitution, an auditable human-factors ontology, a preregistered grounding benchmark, and a population-scale reanalysis of the exposure denominator underlying every rate-based safety metric.

Manuscript7,377 words
References45
Benchmark resultsNone
Peer reviewNone

Read this first: what this document is and is not

This is a working draft. It has not been submitted to any venue, it has not been peer reviewed, and nothing in it should be read as a validated result about any deployed system.

  • The benchmark has no baseline results. Section 5 describes a preregistered apparatus of 68 clause-anchored items. The harness has been exercised only against a mock adapter with a synthetic response profile. No model has been evaluated. Any number attributed to it would be wrong.
  • Four of five internal adversarial reviewers would reject this manuscript as structured - mainly for bundling four separable contributions into one paper, and for presenting an evaluation benchmark with no baselines. The formal-methods reviewer did not reject. That review is reported here rather than buried.
  • The ontology's factor levels are analyst-assigned. The 90 crosswalk alignments are one coder's unadjudicated judgement, with no inter-rater reliability figure. The screening bands are ordinal labels produced by rules marked as convention, never validated against injury or incident outcomes.
  • The simulation studies are synthetic by design. They demonstrate properties of an estimator. They are not findings about workplace safety.
  • The OSHA reanalysis is the one empirical result here. It runs on 2,801,064 real establishment filings and its numbers are reproducible from the committed pipeline. The correction factor it reports is explicitly not a portable constant.
  • Citations. All 45 citation keys used in the text resolve to a reference below, each checked against a primary or publisher-of-record source, with the evidence URL recorded in the repository. 0 entries carry a verification note saying what that check found, including the ones where a page range still needs confirming against printed proceedings.

Cite this

A working draft, not peer reviewed. Cite the manuscript itself, not a finished result.

BibTeX
@misc{the2026,
  author       = {Chimmani, Priyatham},
  title        = {The Hours Denominator: Data Quality in OSHA Establishment Injury Filings and What It Does to Every Benchmark Built on Them},
  year         = {2026},
  howpublished = {\url{https://priyatham9.github.io/grounded/paper-osha.html}},
  note         = {Working draft, not peer reviewed}
}
APA
Chimmani, P. (2026). The Hours Denominator: Data Quality in OSHA Establishment Injury Filings and What It Does to Every Benchmark Built on Them. Working draft, not peer reviewed. https://priyatham9.github.io/grounded/paper-osha.html

Abstract#

Establishment-level injury and illness summaries filed with OSHA's Injury Tracking Application are the only public source from which industry injury-rate benchmarks can be built at scale. We process 2,801,064 establishment filings covering calendar years 2016 to 2024 and show that the hours-worked field, the denominator of every rate computed from these data, is dominated by a small number of implausible entries. Under a plausibility screen of 120 to 4,500 hours per employee per year, 2.07% of filings fail, yet those filings carry 96.7% of all hours reported. The pooled aggregate total recordable incident rate reads 0.134 on the unscreened file and 3.983 after screening. The correction is not a constant: the year-level ratio between screened and unscreened aggregates ranges from 1.39x in 2018 to 249x in 2019, because the share of hours sitting in implausible filings varies from 29% to 99.6% between years. Screened aggregates, by contrast, stay within 3.1 to 4.6 across all 9 years while unscreened aggregates swing between 0.016 and 2.92. Linking establishments across years shows the defect is not random: an establishment flagged one year is flagged the next at an odds ratio of 39.1, against 1.01 under a permutation null that holds each year's flag count fixed, so implausible hours reporting is a property of filers rather than of filings. We characterise the failing records, evaluate the screen's sensitivity to its bounds, examine zero-inflation and overdispersion in recordable-case counts with Poisson, negative binomial and zero-inflated fits, and assess year-over-year stability of peer-group percentile bands. The practical conclusion is narrow and consequential: any benchmark computed from the raw file is wrong by a factor that cannot be known without screening, and no single multiplier repairs it.

1. Introduction#

Every EHS leader eventually asks whether their establishment's injury rate is good. The honest answer requires comparison against establishments of similar size in the same industry, and the data to make that comparison are public: OSHA publishes establishment-level Form 300A summaries each year. They are, however, close to unusable as downloaded. This paper documents why, quantifies the damage, and shows that the damage is not stable enough to correct with a rule of thumb.

The contribution is deliberately modest in scope. We do not propose a new injury model. We show that the denominator underneath every published aggregate is unreliable in a specific, measurable, and year-dependent way, and we publish the reproducible pipeline that establishes it.

2. Data and reproducibility#

Nine public files were processed: the annual ITA data files for CY2016 through CY2022 and the 300A summary extracts covering 2023 through 2025. After de-duplication of 4,703 repeated rows, 2,801,064 filings remained. The pipeline downloads the source files itself, fails loudly if they are unavailable rather than substituting fixtures, and regenerates every table and figure by script. The repository carries 136 tests.

Year-level summary#

YearFilingsFlaggedHours in flaggedTRIR rawTRIR screenedRatio
2016214,9772.30%31.6%2.7293.9211.44x
2017259,7571.66%35.2%2.6343.9981.52x
2018286,8841.56%29.1%2.9204.0691.39x
2019290,4751.60%99.6%0.0164.018249.28x
2020293,3851.91%38.6%2.7924.4911.61x
2021315,9362.60%46.1%2.3854.3681.83x
2022346,7992.02%42.9%2.6614.6051.73x
2023394,2312.34%56.7%1.3803.1472.28x
2024398,6202.37%92.4%0.2803.64212.99x

3 The dependency this paper cannot argue around#

Sections 3 through 5 make a claim about architecture: that a system whose conclusions carry derivation traces back to identified source clauses is auditable in a way that a system conditioned on a persona and a pasted document is not. That claim is about the relationship between a conclusion and a record. It says nothing about the relationship between the record and the world.

The distinction is the one AWS draws explicitly for its own deployed reasoning system. Their documentation states that a VALID result "covers only the parts of the input captured through policy variables," and their own worked example is a claim resting on a forged doctor's note that will be scored valid because no variable captures forgery [aws_arc_docs]. A sound procedure over a false premise returns a sound conclusion about the premise. Nothing about the soundness reaches the world.

For safety-critical question answering in process industries this is not a philosophical footnote, because the records in question are self-reported administrative filings that nothing validates at intake. This section quantifies how bad that is, using the public OSHA Injury Tracking Application (ITA) Form 300A corpus, and then states precisely what follows for a grounded architecture. The short version is that grounding and data validation are separate obligations, and satisfying one does not discharge the other.

All figures in this section were computed by the pipeline in repos/ehs-osha-analysis from nine ITA Form 300A files (CY2016 through CY2024) downloaded from osha.gov, with URLs, byte counts and SHA-256 digests pinned in the repository and verified on 2026-09-03 [osha_ita]. They are transcribed from the generated summary.json and output tables, and regression tests parse the headline values back out of the manuscript's source repository and fail if the prose and the tables disagree. Where a figure comes from the author's earlier work rather than from this pipeline, it is labelled as such.

4 What the corpus is, and what it is not#

OSHA requires annual Form 300A summary submission from establishments with 250 or more employees, and from establishments with 20 to 249 employees in the higher-hazard industry groups listed in 29 CFR 1904 Subpart E [osha_ita_users_guide]. The result is a mandated administrative collection over a selected slice of US industry. It is not a probability sample of the US workforce, and any national rate computed from it is a rate for that slice.

The screened aggregate reported below (about 3.98 recordable cases per 200,000 hours) sits well above the BLS Survey of Occupational Injuries and Illnesses private-industry total recordable case rate of 2.3 per 100 full-time equivalent workers for 2024 [bls2026]. That direction is what a universe skewed toward larger establishments in higher-hazard industries would produce. The two figures are not directly comparable and are not compared here; OSHA publishes its own ITA-versus-SOII comparison and it should be read before anyone treats the gap as a finding [osha_ita_bls_comparison].

Three schema facts matter for anyone building a retrieval or reasoning layer over these files, because each is a silent-wrong-answer trap rather than a loud failure. Column order differs between the 2016–2022 and 2023+ files, and naics_year appears only from

  1. ITA Data CY 2018.csv contains bytes that are not valid UTF-8 (0x92, a Windows-1252 right single quote), so a strict decoder raises and a permissive one silently substitutes. And the size field is not comparable across years: OSHA's summary data dictionary documents codes 1 (<20), 2 (20–249), 21 (20–99), 22 (100–249) and 3 (250+), and states that "code 2 was split to 21 and 22 with the collection of 2023 data" [osha_ita_summary_dict]. The files agree - among plausible filings, code 2 falls from 244,231 in 2022 to 41,343 in 2024 while codes 21 and 22 rise to 224,266 combined - so a pooled panel mixes one 20–249 band with two narrower bands covering the same establishments. The codes are documented; the pooled semantics are not.

None of these is exotic. They are the ordinary condition of regulatory data, and they are invisible to a chunk-and-embed index, which will happily retrieve a 2019 size value and a 2024 size value as if they denoted the same population.

5 The denominator failure#

An incident rate is a ratio:

TRIR = 200,000 × recordable cases / hours worked

The numerator is bounded by how many people work at a site. The denominator is a free-text number on a form. When rates are aggregated across establishments the standard estimator is a ratio of sums, so one filing with an impossible hours value can dominate the denominator of an entire industry, state or national figure while contributing nothing to the numerator. The aggregate is then biased toward zero, and it reads as good news.

Pooled across 2,801,064 deduplicated filings, 2016–2024, under a screen that flags filings outside 120–4,500 hours per average employee per year:

Filings after deduplication2,801,064
Flagged implausible57,857 (2.07%)
Share of all reported hours they hold96.68%
Share of all reported cases they hold1.37%
Aggregate TRIR, unscreened0.134
Aggregate TRIR, screened3.983
Ratio29.7×
0 20 40 60 80 100 % of filings flagged implausible % of all reported hours held by them 2016 2017 2018 2019 2020 2021 2022 2023 2024 A small share of filings carries most of the reported hours Implausible filings hold from a third to nearly all reported hours, year to year. Reporting year Percent
Figure 1. About 2% of filings hold about 97% of all reported hours. The denominator failure is concentrated, not diffuse, which is why a single plausibility bound on hours per employee removes almost all of it.#Source: ehs-osha-analysis/outputs/figures/fig02_hours_share_implausible.svg

The concentration is extreme even within the flagged set. A single filing holds 88.81% of all hours ever reported to the ITA and contributes zero recordable cases: establishment 90427, reporting year 2019, declaring 7 employees and 16,831,620,723,179 hours worked. That is roughly 2.4 trillion hours per employee, against a physical ceiling of 8,760. The ten largest filings by declared hours hold 94.5% of all hours and 0.0008% of all cases. One flag does the entire job: hours per employee > 4,500 fires on 24,617 filings (0.879% of the corpus) carrying 96.65% of all hours. The missing-value and low-hours flags matter for establishment-level rates and are irrelevant to the aggregate.

1 10 100 1000 10000 100000 1e+06 0 1 2 3 4 5 6 7 8 Hours worked per average employee (log10 scale) The tail runs many orders of magnitude beyond anything a workforce can produce (screen: 120-4500 h). log10(annual hours per average employee) Number of filings (log scale)
Figure 2. Most filings sit inside the 120-4,500 hours-per-employee screen bounds; the flagged filings form a right tail that runs many orders of magnitude past the 8,760-hour physical ceiling, out to 10^8 on this axis. The screen removes the tail, not the body.#Source: ehs-osha-analysis/outputs/figures/fig03_hours_per_employee.svg

The screened aggregate is not an artefact of where the bounds are drawn. Across a 25-cell grid of lower bounds (100–400 h) and upper bounds (3,500–6,000 h), the screened aggregate stays within 3.963–3.993. The flagged share is not similarly stable, running from 1.79% to 4.49% across the same grid, and that instability is the point rather than a weakness: the filings that move in and out of the flagged set as bounds shift carry almost none of the hours, so the correction they make is nearly identical. The conclusion survives the threshold choice; the count of flagged filings does not.

6 What does not replicate, and why the non-replication is the finding#

An earlier public repository by the same author ran the same class of plausibility screen over a narrower panel of OSHA establishment filings and reported a headline correction multiplier [chimmani2026]. That multiplier is not reproduced by the pipeline described here, it is superseded by the figures in this section, and it is not restated here as a quantity, because the reason it cannot be reproduced is also the reason it should never have been quoted as one.

The flagged share is stable and does replicate: 2.07% pooled here, and 1.56% to 2.60% across individual years. The multiplier is not a property of the corpus at all. It is a property of whichever years happened to contain an extreme filing:

Table 1. Plausibility screen by reporting year. The flagged share of filings stays near 2% every year while the share of hours flagged and the screened-to-unscreened ratio swing by two orders of magnitude, which is the reason no single correction multiplier is portable.#Source: ehs-osha-analysis/outputs/tables/quality_by_year.csv
labeln_filingsn_implausibleimplausible_sharehours_totalhours_share_implausiblecases_totalcases_share_implausibleaggregate_trir_unscreenedaggregate_trir_screenedratio_screened_to_unscreenedmedian_establishment_trir_screened
201621497749460.02301774630388960.315910571380.017152.7293.9211.4372.791
201725975743230.01664924049175090.351612172000.016072.6343.9981.5172.936
201828688444670.01557936606596950.291513674260.012752.924.0691.3933.001
201929047546500.01601170409485878920.99613735140.012080.016124.018249.32.928
202029338555930.019061019823902890.386214235620.012652.7924.4911.6082.386
202131593682170.026011254982030770.460714967630.012372.3854.3681.8312.533
202234679969930.020161290726462130.429417171800.012512.6614.6051.7312.422
202339423192350.023432257240160540.567415575110.013591.383.1472.282
202439862094330.0236610649440343430.924214931880.015130.28043.64212.991.688
0 1 2 3 4 5 unscreened (all filings) screened 2016 2017 2018 2019 2020 2021 2022 2023 2024 2019 unscreened 0.02 2019 screened 4.02 2024 unscreened 0.28 2024 screened 3.64 Aggregate TRIR before and after the plausibility screen Unscreened TRIR tracks bad hours data, not safety performance. Reporting year Recordable cases per 200,000 hours
Figure 3. The unscreened aggregate TRIR collapses toward zero in 2019 and 2024 because a handful of filings flood the denominator; the screened series stays between roughly 3.1 and 4.6 across all nine years. The gap between the two lines is the correction, and it is a property of the year, not of the corpus.#Source: ehs-osha-analysis/outputs/figures/fig01_aggregate_trir_by_year.svg

The screened rate is stable across nine years (3.15–4.60). The unscreened rate is not (0.016–2.92). The multiplier between them ranges from 1.39× to 249×, and no year subset tried here reproduces the earlier figure. The correct statement of the finding is therefore structural and holds in every single year: a small minority of filings is internally implausible, those filings hold a disproportionate share of the denominator, and removing them raises the aggregate substantially. The specific multiplier is a property of the tail of the hours distribution in whatever slice was taken, and should not be quoted as a constant. We state this here rather than in a limitations paragraph because the earlier framing circulated, and correcting one's own published number in the body of a paper is cheaper than having a reviewer do it.

The methodological point generalizes beyond this corpus. Hanecke and colleagues faced the same problem in 1998 with more than 1.2 million German accident records and no exposure data, and had to construct and compare estimated exposure models before any rate could be computed at all [hanecke1998]. Hopkins reached the same conclusion from the other direction without any statistics, arguing that what determines whether a safety indicator is meaningful is not whether it is labelled leading or lagging but whether, at the level of aggregation in use, there are enough countable events to form a rate - the "zoom effect" [hopkins2009]. Both are arguments about denominators. Neither has an analogue in the performance-shaping-factor literature, which presupposes a well-defined opportunity count throughout [nureg2198], [groth2012].

7 Zero-recordable filings: overdispersion, not a hidden non-reporting population#

Among plausible filings, 37.1% of all establishment-years report zero recordable cases. For NAICS 325 (chemical manufacturing) the pooled figure is 37.4%, rising from 34.1% in 2016 to 41.3% in 2024. The prior repository's often-quoted figure of roughly 38% for chemical establishments is consistent with this panel's 37.4% [chimmani2026].

The obvious reading is that a zero-inflated count model is required - that there exists a distinct population of establishments producing structural zeros through non-reporting. Fitting the models does not support that reading as the main story.

Four intercept-only models with an exposure offset (Poisson, NB2, ZIP, ZINB) were fitted by maximum likelihood and compared by AIC, BIC and observed-versus-expected zeros over the 30 largest NAICS 3-digit groups in 2024, each with at least 500 establishments, covering 306,711 establishment-filings. Industry selection reads only group size, never outcome values or fit statistics.

Table 2. Count-model selection across the 30 largest NAICS 3-digit groups in 2024. Poisson and ZIP never win; NB2 and ZINB split the field under intercept-only fits, so overdispersion is universal and zero inflation is optional; with covariates (text below) ZINB wins in 10 of 30.#Source: ehs-osha-analysis/outputs/tables/count_model_selection_summary.csv
modeln_industries_best_by_aicn_industries_fittedshare_best_by_aic
poisson0300
nb215300.5
zip0300
zinb15300.5
0 2000 4000 6000 8000 10000 0 1 2 3 4 5 6 7 8 9 10+ Recordable-case counts, NAICS 238, 2024: observed vs fitted (size band + NAICS-4 fixed effects, log-hours offset) OSHA ITA Form 300A, reporting year 2024, screened panel, NAICS 238 only; n=20,000 establishments in the fit Recordable cases in the reporting year Number of establishments observed poisson nb2 zinb
Figure 4. For NAICS 238 in 2024, fitted with the covariate-adjusted specification (size band, NAICS-4 fixed effects, log-hours offset) on the 20,000-establishment subsample, Poisson misses the observed zero count badly while NB2 and ZINB track the distribution almost identically; the improvement comes from modelling the spread, not from adding a structural-zero component.#Source: ehs-osha-analysis/outputs/figures/fig04_count_model_fit.svg

Two results point in opposite directions. Overdispersion is universal: Poisson never wins in any of the 30 industries, and the variance-to-mean ratio of recordable counts runs from 3.0 to 756 against the value of 1 that Poisson assumes. ZIP is never selected either, so adding structural zeros to a Poisson does not rescue it - the problem is the spread of the whole distribution rather than the zeros alone. Meanwhile zero-inflation is optional and, where present, small: in 14 of the 30 industries the ZINB inflation parameter collapses to the boundary (on the order of 1e-14) and NB2 wins outright, and where ZINB is selected the inflation probability ranges from 0.001 to 0.069. That is a few percent of establishments, not the 37–41% zero share that motivated fitting it. NB2 frequently predicts slightly more zeros than are observed.

Refitting the four models with a log-hours offset, establishment size-band dummies and NAICS 4-digit fixed effects within each 3-digit group (count_model_covariates_summary.csv in the repository, method in docs/COUNT_MODELS.md) leaves overdispersion intact in all 30 industries (NB2 beats Poisson at boundary-corrected p < 0.001 in every one; NB2 alpha median 0.79, minimum 0.16) but reduces the ZINB wins from 15 of 30 to 10 of 30, with the ZINB-vs-NB2 boundary test below 0.01 in 9. The intercept-only table above therefore overstates zero inflation: part of the excess zeros was mean heterogeneity across size and sub-industry. The conclusion that holds under both specifications is that overdispersion is universal and zero inflation beyond NB2 is present in a minority of industries.

Establishment size accounts for most of the rest. The zero share falls from 0.751 in the 1–19 employee band to about 0.08 above 500 employees, and the median establishment TRIR rises from 0.00 to 3.61 across the same range; a ten-person site has so little exposure that zero is the modal outcome. (The fall is not strictly monotonic - 0.073 at 250–499 rises to 0.082 at 500–999 - and we note it rather than smoothing it.) Any benchmark comparing a site against an industry median without conditioning on size is comparing it against a number driven by how large the other sites are.

0 0.2 0.4 0.6 0.8 1 001-019 020-049 050-099 100-249 250-499 500-999 1000+ Zero-recordable share falls sharply with establishment size A zero rate reflects headcount as much as safety performance. Establishment size band (annual average employees) Share of establishments reporting zero recordables
Figure 5. The zero-recordable share falls from about three quarters in the smallest establishments to under a tenth above 500 employees. Most of the zero share is exposure, which is why size-unconditioned benchmarks compare a site against the wrong population.#Source: ehs-osha-analysis/outputs/figures/fig05_zero_share_by_size.svg

This is a statement about distributional shape, not about reporting behaviour. An overdispersed process and a mixture of compliant and non-compliant reporters can generate similar count distributions, and these data cannot separate them. Nothing here is evidence that under-reporting is absent. What the result does establish is that the bare zero share is not evidence that it is present, which is how that statistic is ordinarily used.

The distributional facts also bound what any downstream predictive layer can do. A variance-to-mean ratio of 241 in the largest size band, combined with a base rate that puts one recordable at roughly one per 10,870 eight-hour worker-shifts at the 2024 BLS private-industry rate [bls2026], is a rare-event regime in which maximum-likelihood logistic regression underestimates event probabilities [king_zeng2001] and in which the standard imbalance corrections - random over- and undersampling, SMOTE [chawla2002] - degrade calibration by strongly overestimating minority-class probability without improving discrimination [goorbergh2022], [carriero2025]. Discrimination metrics are close to uninformative here; calibration is the property that determines whether a score can be acted on [vancalster2019], [vancalster2016].

8 Benchmark bands move on their own#

Peer groups defined as NAICS 3-digit × size band, compared across eight adjacent year pairs over roughly 440 matched cells per pair, are highly reproducible in ordering (median Spearman rho 0.862–0.933 across the p25, p50, p75 and p90 bands) and 95.7–98.5% of publishable cells persist between adjacent years. But the level of each band moves by a median of 9.5–14.3% per year. An establishment sitting exactly on last year's p75 would be roughly a tenth of the way off this year's p75 without any change in its own performance.

0 5 10 15 20 25 0 5 10 15 20 25 Peer-group 75th percentile TRIR: 2023 vs 2024 A benchmark only works if the peer band holds steady between years (469 matched peer groups). p75 TRIR in 2023 p75 TRIR in 2024
Figure 6. Peer-group 75th percentile TRIR values for 2023 and 2024 are well ordered but scatter around the diagonal; the band level moves year to year even when the ranking does not, so an answer at the band edge is close to a coin flip.#Source: ehs-osha-analysis/outputs/figures/fig06_percentile_stability.svg

This bears directly on the question-answering task. A retrieval system asked "is this site's TRIR above the industry 75th percentile?" will return an answer that is approximately a coin flip near the band edge, and it will return it with the same fluency whether the site is clearly above, clearly below, or inside the year-to-year noise. The band is a real quantity with a real sampling distribution; the answer is presented as a fact.

9 Hour-of-shift: a claim withdrawn, and the conditions for ever making it again#

The author's earlier work circulated a shift-timing claim about when during a shift injuries peak [chimmani2026]. It is withdrawn here rather than restated, and its numbers are not reproduced anywhere in this manuscript, for a reason that is structural rather than a matter of degree: the pipeline described in this section could not produce such a figure at all. Form 300A is an annual summary. It carries case counts and total hours worked and no time-of-day, time-employee-began-work, or narrative field at all [osha_ita_summary_dict]. Hour-of-shift requires Form 300/301 case-detail data, which OSHA began publishing on a different and narrower establishment universe with the 2023/2024 collection cycle [osha_ita_case_detail_dict_2026]. Any future hour-of-shift figure must name its own dataset, year range and N, must not be read as a subset of the 300A panel analysed here, and must not inherit this panel's filing count.

Two further constraints apply, and both are decisive for how the finding may be phrased.

First, the phenomenon is named and roughly twenty-five years old. Tucker, Sytnik, Macdonald and Folkard called it the "2–4 h shift phenomenon" [tucker2000], and Folkard and Tucker report "a slightly heightened risk from the second to the fifth hour" in their review of shift work, safety and productivity [folkard2003]. A large-scale corroboration in a new sector would be a real contribution. A discovery claim is not available, and no corroboration is offered in this manuscript.

Second, and more seriously, a peak in the count distribution is not a peak in risk. Any statistic of the form "such-and-such a share of injuries occurs in the first four hours" is by construction a share of counts, and counts are governed by how many people are at work in each hour of shift. Folkard and Tucker found that once exposure is accounted for, risk rises approximately exponentially with time on shift, more than doubling by the twelfth hour relative to the first eight [folkard2003] - a pattern that an unadjusted count distribution, dominated by the hours in which most people are working, will not show. Hanecke and colleagues had to build estimated exposure models precisely because the hours-at-work denominator was unavailable in the German data [hanecke1998], and no public US dataset supplies an hours-at-risk-by-shift-hour denominator either. MSHA's Accidents file carries SHIFT_BEGIN_TIME and an accident time on essentially all records [msha_accidents_definition], and BLS publishes an "hours worked before event" dimension in its case-characteristics series [bls_ca_documentation], so cross-source comparison of the distribution is feasible. The exposure denominator is not.

The defensible statement is therefore about the observed distribution of reported incidents, explicitly not about risk, with the denominator gap stated as an open problem rather than assumed away. We adopt that phrasing throughout and recommend it to anyone reusing the figure.

10 How the two failures compose#

The preceding subsections describe defects in a data corpus. Sections 4 and 5 describe defects in ungrounded language models: parametric recall that degrades on long-tail content [kandpal2023], [mallen2023], training and evaluation regimes that reward a plausible guess over an abstention [kalai2025], and distributional representations that encode relatedness without encoding relation type, so that taxonomic siblings are the maximal-confusion class. These are independent failure modes with independent causes. They compose in three specific ways, and the composition is worse than either alone.

A retrieved number is not a validated number. Retrieval-augmented generation supplies non-parametric memory and marginalizes over retrieved documents [lewis2020]. That machinery grounds an assertion in a document. It carries no commitment about whether the document is internally consistent, and none of the standard RAG evaluation metrics test for it: attribution frameworks score whether a claim is attributable to its cited source [rashkin2023], which a filing declaring 2.4 trillion hours per employee satisfies perfectly. A system that retrieves establishment 90427's filing and reports a TRIR of 0.00 has produced a fully attributable, fully traceable, entirely wrong answer. The derivation trace this paper advocates certifies that the conclusion follows from the record. Whether the record is admissible is a separate predicate, and it has to be computed.

Aggregation hides the defect from the reader and from the model. The single filing above is visible at establishment level - 7 employees, 16.8 trillion hours - and invisible in an industry aggregate, where it appears only as a suspiciously low rate. The failure mode is therefore worst at exactly the level of aggregation at which executives and benchmarking tools operate. A model asked for an industry rate has no signal in the retrieved aggregate telling it the aggregate is broken, and the resulting answer is low-variance across paraphrase and across sampling, so consistency-based hallucination detectors [manakul2023] and semantic-entropy methods [kuhn2023], [farquhar2024] will score it confident. Uncertainty estimation detects the model's uncertainty. It does not detect the corpus's.

Grounding narrows the error class without eliminating it, and the residual is measurable. The governing empirical result is Magesh and colleagues' preregistered evaluation of commercial legal research tools marketed as hallucination-free, which measured hallucination rates between 17% and 33% [magesh2025]. There is no basis for expecting a process-safety deployment to do better, and every reason - hierarchical, cross-referenced, edition-versioned, table-heavy source documents - to expect it to face harder retrieval conditions. Any claim in this paper is therefore a measured delta under stated conditions, never an elimination.

The constructive consequence is that plausibility screening belongs inside the architecture rather than upstream of it, as a checkable predicate over each retrieved record. The screen described in §6.3 is six deterministic predicates over declared hours, declared employees and case components; we call it that, and not automated reasoning, because it has no decision procedure over a formal semantics, no complexity characterization, and emits no proof object. Stating a filing's admissibility as a constraint problem - given declared hours, employee count and case counts, is this record consistent with the physical and regulatory constraints? - would produce an artifact that a formal-methods reviewer would recognize, namely an unsat core naming which constraints a specific filing violates [barrett2021smt], [demoura2008]. We have not built that, and we flag it as future work rather than as a contribution. The precedent for formalizing a body of regulation as executable logic, with the formalization kept traceable to the statutory text clause by clause, is forty years old [sergot1986]; we found no OSHA successor to it.

A closing note on verification that checks the wrong invariant. The benchmark corpus described in Section 7 originally reported 68 of 68 items verified against live eCFR text with zero failures. An independent audit found that one item cited 29 CFR 1910.157(d) while its anchor text actually appears in paragraph (e)(2): the verifier confirmed only that the cited paragraph markers appeared somewhere in the retrieved section, so a clause-to-anchor mismatch was structurally invisible to it. The headline "68/68, 0 failures" was true and weaker evidence than it looked. The defect was repaired, an anchor-locality check was added, and the corpus re-verified. We report this because it is the same failure as the OSHA denominator, one level up: a validation procedure that returns a clean result on a broken record, because the property it tests is not the property that matters. Leakage between predictor construction and outcome labels has the same shape and is documented inside safety machine learning specifically - Baker, Hallowell and Tixier rebuilt an earlier construction-injury prediction study with independent human annotation to eliminate artificial correlation between predictors and predictands [baker2020ai] - and across 294 papers in 17 fields more broadly [kapoor2023]. In each case the pipeline ran, the tests passed, and the number was wrong.

The practical requirement that follows is unglamorous and is the precondition for everything else in this paper: state the dataset, the year range, the N, the screen, and the denominator, separately, for every number. Report calibration rather than discrimination where a score will trigger an action [steyerberg2010], [collins2015tripod]. And treat the admissibility of a retrieved record as something the system computes and exhibits, rather than something the retriever's confidence score is assumed to have covered.


11 Why the hours fail: implausible reporting is a property of establishments, not of filings#

Sections 5 and 6 establish that the exposure denominator fails, and that the failure is not stable enough to correct with a constant. Neither explains why. Two mechanisms predict opposite things. If implausible hours are keying errors, they should strike establishments independently from one year to the next, and knowing that an establishment filed implausibly in one year should tell you nothing about the next. If instead certain establishments report hours in the wrong unit as a matter of routine, because of how their filing process or payroll export is configured, bad filings should concentrate in the same establishments year after year.

The Injury Tracking Application carries a stable establishment_id, so this is directly testable. Linking the 1,238,236 distinct establishments across CY2016 to CY2024 yields 614,928 that appear in two or more years and 16,307 present in all nine.

Conditioning on consecutive-year pairs, an establishment flagged in one year is flagged again the next with probability 27.2% , against 0.95% for establishments not previously flagged. The pooled persistence odds ratio is 39.1, and every one of the eight year-pairs lies between 24.7 and 63.1.

An odds ratio is only meaningful against a null. Holding each year's number of flagged filings fixed and reassigning them uniformly at random across the establishments filing that year, over 200 permutations at seed 20260909, gives a mean odds ratio of 1.012 with a 95% interval of [0.924, 1.092]. The observed value of 39.1 lies far outside that interval, and 0% of permutations reach it.

The 2019 file, where 1.60% of filers are flagged and almost all reported hours sit in flagged records, follows the same pattern rather than standing apart. Establishments flagged in 2019 and also present in 2018 were flagged in 2018 29.3% of the time against a 2018 baseline of 1.56%, an enrichment of 18.8x; the corresponding 2020 enrichment is 16.4x.

0 20 40 60 2016-2017 2017-2018 2018-2019 2019-2020 2020-2021 2021-2022 2022-2023 2023-2024 pooled permutation null Flag-persistence odds ratio: observed vs. permutation null 200 permutations, seed 20260909; null holds each year's flag count fixed Year pair Odds ratio year-pair pooled observed permutation mean
Figure 7. Persistence of the implausibility flag across consecutive filing years. The observed odds ratio sits far above a permutation null in which the same number of flags is reassigned at random each year, so implausible hours reporting attaches to establishments rather than arriving at random.#Source: ehs-osha-analysis/outputs/figures/fig08_panel_persistence.svg

Two qualifications matter. Cohen's kappa for the pooled transition is 0.28, which is fair rather than strong agreement, so a persistent subset coexists with genuinely one-off errors rather than replacing them. And while establishment_id is populated on every row and stable in state for 99.7% of identifiers, 313,527 of them (25.3%) carry more than one EIN across years and 108,855 (8.8%) more than one company name, so the linkage is good but not perfect and some persistence may reflect corporate restructuring rather than a single continuing filer.

The practical consequence is narrow and useful. Because the defect attaches to establishments and persists, a screen is not the only available remedy: an establishment with a consistent unit error across years is in principle repairable rather than merely excludable, and a filer-level correction is a more promising direction than any single global multiplier of the kind Section 6 rules out.

Limitations#

This paper contributes four things: a structural account of why ungrounded retrieval substitutes semantically adjacent answers in safety-critical technical question answering; an auditable encoding of a performance-influencing-factor vocabulary with an explicit crosswalk to the established frameworks it is drawn from; a rule engine whose screening output carries a derivation trace; and an open, source-verified benchmark corpus for measuring grounding.

Three of the four are instruments rather than results. The benchmark has not been run against any real system. The ontology has never been fitted to outcome data. The structural-equation work is a set of simulation studies of an estimator, not an analysis of injuries. Only the OSHA denominator analysis reports empirical findings, and its findings are about the integrity of a reported quantity rather than about safety performance.

The corpus holds 68 items across eight regulatory domains, of which 63 are factual and 5 rest on a false premise, organized into 23 complete minimal pairs and 44 question families. Every item's citation was re-checked against the eCFR versioner API, and the committed verification report records 68 of 68 verified with zero failures across 19 CFR sections. That number establishes provenance and nothing else. It says the cited paragraph exists and contains the quoted anchor. It does not say the answer key is a good answer.

Scoring is lexical, and that has a price. A paraphrase that avoids every listed surface form of a required concept is scored as a miss. The choice was made because the benchmark's subject is whether a system distinguishes near-synonymous technical entities, and scoring that with an embedding model would use as the measuring instrument the very mechanism under study. But the contrast heuristic that separates "a rupture disk has no blowdown" from an assertion of blowdown operates on a fixed character window and misfires in both directions. Any adjacent-substitution rate reported from this harness is a screen. Automatic attribution evaluation is itself unreliable [yue2023], which is why the established practice is a human-adjudicated subsample with inter-rater agreement reported [rashkin2023, honovich2022true]. No human adjudication has been performed here.

The four contextual categories encoded here are adopted, not invented. IDHEAS-G organizes 20 performance-influencing factors into exactly four context categories, and states directly that "PIFs are classified according to the four types of context: environment and situation, system, personnel, and task" [nureg2198]. Task and System are identical labels; Operational Context and Human Context are synonyms for Environment and Situation, and Personnel. The correspondence is one to one, and it maps further onto CREAM's nine common performance conditions [hollnagel1998], SPAR-H's eight PSFs [gertman2005], NUREG-1792's fifteen, and HSE's job, person and organisation headings [hse_pifs]. Presenting this vocabulary as novel would be indefensible, and we do not. What is contributed is the encoding, the crosswalk with per-row citations and asserted absences, and the traceability. That is a smaller claim, and it is the one the artifact supports.

A latent score is not a probability. It has no link function, no exposure denominator and no time window, so no threshold on it can be justified. Separately, factor scores are not uniquely determined by a fitted model: on the worked example, two equally valid sets of scores for the same factor correlate as low as 0.54 to 0.59 [steiger1979, grice2001]. A per-crew or per-shift score inherits that indeterminacy however well the model fits.

References#

aws_arc_docs
Amazon Web Services, "Automated Reasoning checks concepts," Amazon Bedrock User Guide. https://docs.aws.amazon.com/bedrock/latest/userguide/automated-reasoning-checks-concepts.html evidence
baker2020ai
Baker, H., Hallowell, M. R., & Tixier, A. J.-P. (2020). AI-based prediction of independent construction safety outcomes from universal attributes. Automation in Construction, 118, 103146. https://doi.org/10.1016/j.autcon.2020.103146 evidence
barrett2021smt
Clark Barrett, Roberto Sebastiani, Sanjit A. Seshia and Cesare Tinelli, "Chapter 33: Satisfiability Modulo Theories," in Handbook of Satisfiability, 2nd edition, Frontiers in Artificial Intelligence and Applications, IOS Press, 2021. DOI 10.3233/faia201017 evidence
bls2026
U.S. Bureau of Labor Statistics (2026). Employer-Reported Workplace Injuries and Illnesses - 2023–2024. News release USDL-26-0101, January 22, 2026. evidence
bls_ca_documentation
U.S. Bureau of Labor Statistics. Occupational Injuries and Illnesses - Characteristics Data (ca): ca.txt survey documentation. https://download.bls.gov/pub/time.series/ca/ca.txt evidence
carriero2025
Carriero, A., Luijken, K., de Hond, A., Moons, K. G. M., van Calster, B., & van Smeden, M. (2025). The Harms of Class Imbalance Corrections for Machine Learning Based Prediction Models: A Simulation Study. Statistics in Medicine, 44(3–4), e10320. https://doi.org/10.1002/sim.10320 evidence
chawla2002
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953 evidence
chimmani2026
Priyatham Chimmani, ehs-benchmarks: open OSHA injury-rate benchmarks, GitHub repository, MIT licence, 2026. https://github.com/priyatham9/ehs-benchmarks evidence
collins2015tripod
Collins, G. S., Reitsma, J. B., Altman, D. G., & Moons, K. G. M. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD Statement. BMC Medicine, 13, 1. https://doi.org/10.1186/s12916-014-0241-z evidence
demoura2008
Leonardo de Moura and Nikolaj Bjørner, "Z3: An Efficient SMT Solver," TACAS 2008, Lecture Notes in Computer Science 4963, pp. 337–340, 2008. DOI 10.1007/978-3-540-78800-3_24 evidence
farquhar2024
Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017), 625-630. DOI: 10.1038/s41586-024-07421-0 evidence
folkard2003
Folkard, S. & Tucker, P. (2003). Shift work, safety and productivity. Occupational Medicine, 53(2), 95-101. doi:10.1093/occmed/kqg047 evidence
gertman2005
Gertman, D.I., Blackman, H.S., Marble, J.L., Byers, J.C. & Smith, C.L. (2005). The SPAR-H Human Reliability Analysis Method. NUREG/CR-6883, INL/EXT-05-00509. Washington, DC: U.S. Nuclear Regulatory Commission. Manuscript completed September 2004; published August 2005. evidence
goorbergh2022
van den Goorbergh, R., van Smeden, M., Timmerman, D., & Van Calster, B. (2022). The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression. Journal of the American Medical Informatics Association, 29(9), 1525–1534. https://doi.org/10.1093/jamia/ocac093 evidence
groth2012
Groth, K.M. & Mosleh, A. (2012). A data-informed PIF hierarchy for model-based Human Reliability Analysis. Reliability Engineering & System Safety, 108, 154-174. evidence
hanecke1998
Hanecke, K., Tiedemann, S., Nachreiner, F. & Grzech-Sukalo, H. (1998). Accident risk as a function of hour at work and time of day as determined from accident data and exposure models for the German working population. Scandinavian Journal of Work, Environment & Health, 24(Suppl 3), 43-48. PMID 9916816. evidence
hollnagel1998
Hollnagel, E. (1998). Cognitive Reliability and Error Analysis Method (CREAM). Oxford: Elsevier Science. evidence
hopkins2009
Hopkins, A. (2009). Thinking about process safety indicators. Safety Science, 47(4), 460–465. https://doi.org/10.1016/j.ssci.2007.12.006 evidence
hse_pifs
Health and Safety Executive (n.d.). Performance Influencing Factors (PIFs). HSE Human Factors guidance. https://www.hse.gov.uk/humanfactors/assets/docs/pifs.pdf evidence
kalai2025
Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664 evidence
kandpal2023
Kandpal, N., Deng, H., Roberts, A., Wallace, E., & Raffel, C. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. Proceedings of ICML 2023. arXiv:2211.08411 evidence
kapoor2023
Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804 evidence
king_zeng2001
King, G., & Zeng, L. (2001). Logistic Regression in Rare Events Data. Political Analysis, 9(2), 137–163. https://doi.org/10.1093/oxfordjournals.pan.a004868 evidence
kuhn2023
Kuhn, L., Gal, Y., & Farquhar, S. (2023). Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation. Proceedings of ICLR 2023 (Spotlight). arXiv:2302.09664 evidence
lewis2020
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arXiv:2005.11401 evidence
magesh2025
Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies, 22, 216-242. DOI: 10.1111/jels.12413 evidence
mallen2023
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023). When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. Proceedings of ACL 2023. arXiv:2212.10511 evidence
manakul2023
Manakul, P., Liusie, A., & Gales, M. J. F. (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. Proceedings of EMNLP 2023. arXiv:2303.08896 evidence
msha_accidents_definition
Mine Safety and Health Administration. Accidents Definition File, MSHA Open Government Data. https://arlweb.msha.gov/OpenGovernmentData/DataSets/Accidents_Definition_File.txt evidence
nureg2198
U.S. Nuclear Regulatory Commission (2021). The General Methodology of an Integrated Human Event Analysis System (IDHEAS-G). NUREG-2198. Washington, DC: Office of Nuclear Regulatory Research. Manuscript completed November 2020; published May 2021. evidence
osha_ita
Occupational Safety and Health Administration. Establishment-Specific Injury and Illness Data (Injury Tracking Application). U.S. Department of Labor. https://www.osha.gov/Establishment-Specific-Injury-and-Illness-Data evidence
osha_ita_bls_comparison
Occupational Safety and Health Administration. Comparison Between OSHA ITA Data and BLS SOII Estimates. https://www.osha.gov/sites/default/files/ComparisonBetweenOSHAITA_Data_and_BLS_SOII_Estimates.pdf evidence
osha_ita_case_detail_dict_2026
Occupational Safety and Health Administration. ITA Case Detail Data Dictionary (2026 revision). https://www.osha.gov/sites/default/files/case_detail_data_dictionary_2026.pdf evidence
osha_ita_summary_dict
Occupational Safety and Health Administration. ITA Summary Data Dictionary (Form 300A). https://www.osha.gov/sites/default/files/summary_data_dictionary.pdf evidence
osha_ita_users_guide
Occupational Safety and Health Administration. Injury Tracking Application (ITA) Data Users Guide. Last updated January 2025. https://www.osha.gov/sites/default/files/ITA_data_users_guide.pdf evidence
rashkin2023
Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., & Reitter, D. (2023). Measuring Attribution in Natural Language Generation Models. Computational Linguistics, 49(4), 777-840. DOI: 10.1162/coli_a_00486 evidence
sergot1986
M. J. Sergot, F. Sadri, R. A. Kowalski, F. Kriwaczek, P. Hammond and H. T. Cory, "The British Nationality Act as a logic program," Communications of the ACM 29(5):370–386, 1986. DOI 10.1145/5689.5920 evidence
steyerberg2010
Steyerberg, E. W., Vickers, A. J., Cook, N. R., Gerds, T., Gonen, M., Obuchowski, N., Pencina, M. J., & Kattan, M. W. (2010). Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology, 21(1), 128–138. https://doi.org/10.1097/EDE.0b013e3181c30fb2 evidence
tucker2000
Tucker, P., Sytnik, N., Macdonald, I. & Folkard, S. (2000). Temporal determinants of accident risk: the '2-4 h shift phenomenon'. In: Hornberger, S., Knauth, P., Costa, G. & Folkard, S. (eds.), Shiftwork in the 21st Century. Frankfurt: Peter Lang, 99-105. evidence
vancalster2016
Van Calster, B., Nieboer, D., Vergouwe, Y., De Cock, B., Pencina, M. J., & Steyerberg, E. W. (2016). A calibration hierarchy for risk models was defined: from utopia to empirical data. Journal of Clinical Epidemiology, 74, 167–176. https://doi.org/10.1016/j.jclinepi.2015.12.005 evidence
vancalster2019
Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W., on behalf of Topic Group 'Evaluating diagnostic tests and prediction models' of the STRATOS initiative (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17, 230. https://doi.org/10.1186/s12916-019-1466-7 evidence
yue2023
Yue, X., Wang, B., Chen, Z., Zhang, K., Su, Y., & Sun, H. (2023). Automatic Evaluation of Attribution by Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2023, 4615-4635. DOI: 10.18653/v1/2023.findings-emnlp.307 evidence
grice2001
James W. Grice, "Computing and evaluating factor scores," Psychological Methods 6(4), 2001, pp. 430-450. DOI: 10.1037/1082-989X.6.4.430. evidence
honovich2022true
Honovich, O., Aharoni, R., Herzig, J., Taitelbaum, H., Kukliansy, D., Cohen, V., Scialom, T., Szpektor, I., Hassidim, A., & Matias, Y. (2022). TRUE: Re-evaluating Factual Consistency Evaluation. Proceedings of the Second DialDoc Workshop, 161-175. DOI: 10.18653/v1/2022.dialdoc-1.19 evidence
steiger1979
James H. Steiger, "Factor indeterminacy in the 1930's and the 1970's: some interesting parallels," Psychometrika 44(1), 1979, pp. 157-167. DOI: 10.1007/BF02293967. evidence