Programme metrics · every figure read from repo outputs
Observatory.
One screen for the whole programme. Hover, focus or tap any mark for its exact value; every chart has a data table and names the file it came from.
Status: draft research, not peer reviewed. The grounding benchmark shows baselines only: no language model evaluated.
The screen barely moves the rate
25 hours-per-employee plausibility windows. The share of filings flagged changes more than twofold across the grid; the screened aggregate TRIR stays within a narrow range.
One year, one broken denominator
Aggregate TRIR by filing year before and after screening. The 2019 file carries implausible reported hours that pull the unscreened rate toward zero. Switch to log scale to see how far.
Where implausible filings sit
Share of filings flagged implausible, by two-digit NAICS sector (sectors with at least 1,000 filings) and by filing year.
Peer bands, year over year
Establishment TRIR percentiles for a NAICS-3 by size-band peer group (publishable groups only), and how stable each percentile's ranking is between consecutive years.
Flags persist
Among establishments filing in consecutive years, how often an implausibility flag in one year is followed by another the next year.
Which count model fits the zeros
Observed zero-case establishments divided by the zeros each model expects, for 30 NAICS-3 industries in 2024. 1.0 means the model reproduces the zeros exactly. A ring marks the best model by AIC.
Grounding benchmark baselines
Accuracy with intervals for the reference adapters. These bracket the task; no language model evaluated.
SEM recovery by sample size
Simulation study 01: coverage of nominal 95% analytic intervals for each structural path, and the ratio of analytic to empirical standard error.
Ontology crosswalk coverage
Each ontology factor against four external human factors frameworks. Cells show the strongest recorded match; hover for the external concepts.