Where the method breaks
A structural equation model can fit a safety dataset perfectly and still return a number that means something other than what the slide says it means. These six scenes show exactly where that happens, using simulations where the true answer is known in advance.
All data on this page is synthetic. Every dataset was generated from a known model so that the estimate can be compared with the truth. No plant, contractor or worker data was used. Nothing here is a finding about a real site.
The model under review
On average the method is right (bias 0.007), but one study can miss by 0.152
The estimator recovers the coefficients. A single study still lands anywhere inside a band the slide never shows.
On average the method is right (bias 0.007), but one study can miss by 0.152
Four predictors generate the data at 0.45, 0.30, -0.25 and -0.20
Unsafe acts, operational stress, system condition and safety response capability each load on a latent and each points at a latent risk score. Those four generating values are known here and hidden from the estimator, which is the only reason the rest of this page is measurable.
Largest bias across the four paths is 0.007
That is at n = 100; at n = 5000 the largest bias is 0.0014. Across 300 replications there were no improper solutions and no failed fits at any sample size. Whatever goes wrong later is not bias in the coefficients.
At n = 100 a single study lands between 0.15 and 0.76
The empirical standard deviation on the 0.45 path is 0.152 at n = 100 and 0.041 at n = 1000. Drag the slider to watch the band close. Your report is one draw from that band, and the slide only ever shows the draw.
The textbook 95 percent interval holds only 82 times in 100
Every range your team reports from the analytic formula is about a third too narrow, and a larger sample does not close it.
The textbook 95 percent interval holds only 82 times in 100
A 95 percent interval promises 95 hits in 100
That promise is the whole reason a coefficient is reported with a range instead of alone.
Measured coverage runs 0.757 to 0.863 at every sample size
Across 300 replications the analytic interval never climbs toward 0.95 as n grows. The analytic standard error is about two thirds of the real spread: the ratio is 0.63 at n = 100 and 0.72 at n = 5000. More data does not fix it, because the error is in what the formula conditions on.
The bootstrap covers 0.925 to 0.975 at n = 500
The analytic interval covers 0.775 to 0.85 on the same data. The bootstrap interval is about 40 percent wider. Ask for the wider one: it reports the precision you actually had.
A passing fit needs 153 units. Naming the bigger driver needs 6,884.
Most safety models are sized for the cheapest claim and then used to make the most expensive one.
A passing fit needs 153 units. Naming the bigger driver needs 6,884.
153 units buy a fit statistic
Eighty percent power for the test of close fit at df = 80 needs 153 units. Most study designs are sized against this number, and it is the cheapest claim on the list.
369 units buy a usable coefficient, 1,459 a sharp one
A standard error of 0.05 on one standardized path needs 369 units, and halving it to 0.025 needs 1,459. Precision costs four times as much every time you halve it, so a sample sized for fit cannot carry a coefficient claim.
6,884 units before you can name the bigger driver
Separating 0.45 from 0.30 at 80 percent power takes 770 units; separating -0.25 from -0.20 takes 6,884. "Unsafe acts matter more than system condition" is a different claim from "both paths are non-zero" and costs 45 times the data. Before you rank drivers in a steering meeting, get the n.
Leave out one common cause and a true 0.20 reads 0.495, with a perfect fit
Four models with the wrong causal structure all clear the fit cutoffs, and one of them inflates a coefficient by 0.295.
Leave out one common cause and a true 0.20 reads 0.495, with a perfect fit
Drop one common cause and 0.173 becomes 0.495
The true path is 0.20. With the confounder in the model it is estimated at 0.173; drop the confounder and the same path reads 0.495, a bias of +0.295. The trimmed model has CFI 1.000 and RMSEA 0.000, so it fits better than the correct one.
Reverse the arrow and the fit is CFI 0.9999
Data generated as injuries driving climate, fitted the other way round, returns -0.570 with p below 0.001 and RMSEA 0.004. Cause to effect and effect to cause give identical output: beta 0.4927, SE 0.0159, chi-square 7.696 on 8 df.
All four wrong models clear the conventional cutoffs
Omitting a mediator moves the same coefficient from the direct effect 0.182 to the total effect 0.442, and the fit still passes. A fit index only says the implied covariance matrix is close to the observed one, and several causal structures imply the same matrix. A good fit is not evidence that the structure is right.
With rare events, 680 alerts for every real one
At a base rate of 9.2 events per 100,000 worker-shifts, the metric on the scoreboard and the experience of the supervisor stop agreeing.
With rare events, 680 alerts for every real one
680 alerts for every one real event
That is a model with 80 percent sensitivity and 95 percent specificity at a recordable rate of 9.2 events per 100,000 worker-shifts. Push specificity to 99.99 percent and the burden falls to 2.1 alerts per event.
AUC moves 0.007 while alerts per event move 39-fold
The same generating signal gives AUC 0.685 at a 30 percent base rate and 0.692 at 0.5 percent. Precision in the top five percent falls from 0.629 to 0.016 and alerts per true event climb from 1.6 to 62. AUC cannot see the base rate your sites actually have.
Undersampling overstates risk by a factor of 88
Trained as collected, AUC is 0.711 and predicted risk matches observed to within two percent. Undersample to a 50 percent event rate and AUC is 0.707, unchanged, while the Brier score goes from 0.005 to 0.230. Sizing is the same story: ten percent relative precision needs 384 events, or 4,175,499 worker-shifts. Ask any alerting vendor for alerts per true event, not AUC.
The same fitted coefficients can mean an effect of 0.45 or of 0.281
Three causal structures return the same fitted numbers and three different answers to what happens if you act.
The same fitted coefficients can mean an effect of 0.45 or of 0.281
Three worlds return the same four coefficients
The fitted vector is [0.448, 0.306, -0.254, -0.200] in world one, [0.453, 0.307, -0.260, -0.208] in world two and [0.448, 0.303, -0.246, -0.203] in world three. Rounded for a slide, all three are the same slide.
The effect of acting is 0.45 in world one and 0.281 in world three
Only in world one is the coefficient on unsafe acts the effect of intervening on unsafe acts. In world two the reported -0.260 on system condition is a direct effect and the total effect is -0.070. In world three an unmeasured common cause sits behind unsafe acts and risk, so intervening returns 0.281 while the regression still prints 0.448.
4.6 percent of units cross the audit flag when the scale changes
Crew A scores 2.55 and crew B 3.15 on the same raw inputs at both sites, yet standardizing within site puts crew A ahead at site 1 and crew B ahead at site 2, because overtime varies more at one of them. Fit the measurement model by site and one loading drops from 0.769 to 0.301 and composite reliability from 0.823 to 0.675. Forcing site A's model onto site B sends audits to a different 4.6 percent of units.
What this does not establish
Every dataset on this page is synthetic, generated from a known model by simulations/dgp.py. That is the design, not a shortcut: the failures shown here are only visible when the true answer is known. Nothing here describes a real site, a real workforce or a real rate.
These are simulation studies of estimator behaviour. They do not show that any particular published safety model is wrong. They show which claims a model of this shape can and cannot support at a given sample size.
The recovery study uses 300 replications per cell; the bootstrap study uses 40, with a Monte Carlo standard error on coverage of about 0.034. Read the bootstrap coverage figures as approximately 0.95, not exactly.
No manuscript from this work has been submitted or peer reviewed. The code, the generating models and the result tables are all in the repository so the numbers can be regenerated and disagreed with.
One thing to do on Monday
Take the last risk model your team presented. Ask for the sample size and the interval method behind the coefficient that drove the decision. If the interval came from the analytic standard error, ask for a bootstrap interval instead, and ask what would have to be true about the causal structure for that coefficient to be the effect of acting on it.
- 1 Get the n. Compare it against 369 for a usable path and 6884 for a claim that one path beats another.
- 2 Replace the analytic interval with a bootstrap interval. Expect it to be about 40 percent wider.
- 3 Write down the one causal structure under which the coefficient answers the question being asked. If you cannot, report it as an association.
- 4 For any alerting model, ask for alerts per true event at the operating point, not AUC.