Where the method breaks

A structural equation model can fit a safety dataset perfectly and still return a number that means something other than what the slide says it means. These six scenes show exactly where that happens, using simulations where the true answer is known in advance.

All data on this page is synthetic. Every dataset was generated from a known model so that the estimate can be compared with the truth. No plant, contractor or worker data was used. Nothing here is a finding about a real site.

The model under review

On average the method is right (bias 0.007), but one study can miss by 0.152

The estimator recovers the coefficients. A single study still lands anywhere inside a band the slide never shows.

Scene 01

On average the method is right (bias 0.007), but one study can miss by 0.152

300 replications per point
Source: results/study01_recovery.csv. Bars show the mean estimate; the band is one empirical standard deviation across replications.
The model

Four predictors generate the data at 0.45, 0.30, -0.25 and -0.20

Unsafe acts, operational stress, system condition and safety response capability each load on a latent and each points at a latent risk score. Those four generating values are known here and hidden from the estimator, which is the only reason the rest of this page is measurable.

The estimates

Largest bias across the four paths is 0.007

That is at n = 100; at n = 5000 the largest bias is 0.0014. Across 300 replications there were no improper solutions and no failed fits at any sample size. Whatever goes wrong later is not bias in the coefficients.

The spread

At n = 100 a single study lands between 0.15 and 0.76

The empirical standard deviation on the 0.45 path is 0.152 at n = 100 and 0.041 at n = 1000. Drag the slider to watch the band close. Your report is one draw from that band, and the slide only ever shows the draw.

The textbook 95 percent interval holds only 82 times in 100

Every range your team reports from the analytic formula is about a third too narrow, and a larger sample does not close it.

Scene 02

The textbook 95 percent interval holds only 82 times in 100

Source: results/study01_recovery.csv (analytic) and results/study01_bootstrap.csv (bootstrap, n = 500, 40 replications, Monte Carlo SE about 0.034).
The claim

A 95 percent interval promises 95 hits in 100

That promise is the whole reason a coefficient is reported with a range instead of alone.

The measurement

Measured coverage runs 0.757 to 0.863 at every sample size

Across 300 replications the analytic interval never climbs toward 0.95 as n grows. The analytic standard error is about two thirds of the real spread: the ratio is 0.63 at n = 100 and 0.72 at n = 5000. More data does not fix it, because the error is in what the formula conditions on.

The fix

The bootstrap covers 0.925 to 0.975 at n = 500

The analytic interval covers 0.775 to 0.85 on the same data. The bootstrap interval is about 40 percent wider. Ask for the wider one: it reports the precision you actually had.

A passing fit needs 153 units. Naming the bigger driver needs 6,884.

Most safety models are sized for the cheapest claim and then used to make the most expensive one.

Scene 03

A passing fit needs 153 units. Naming the bigger driver needs 6,884.

Source: results/study01_requirements.csv and results/study01_distinguishability.csv. Log scale.
Fit

153 units buy a fit statistic

Eighty percent power for the test of close fit at df = 80 needs 153 units. Most study designs are sized against this number, and it is the cheapest claim on the list.

Precision

369 units buy a usable coefficient, 1,459 a sharp one

A standard error of 0.05 on one standardized path needs 369 units, and halving it to 0.025 needs 1,459. Precision costs four times as much every time you halve it, so a sample sized for fit cannot carry a coefficient claim.

Ranking

6,884 units before you can name the bigger driver

Separating 0.45 from 0.30 at 80 percent power takes 770 units; separating -0.25 from -0.20 takes 6,884. "Unsafe acts matter more than system condition" is a different claim from "both paths are non-zero" and costs 45 times the data. Before you rank drivers in a steering meeting, get the n.

Leave out one common cause and a true 0.20 reads 0.495, with a perfect fit

Four models with the wrong causal structure all clear the fit cutoffs, and one of them inflates a coefficient by 0.295.

Scene 04

Leave out one common cause and a true 0.20 reads 0.495, with a perfect fit

Source: results/study03_confounder.csv, study03_reverse.csv, study03_equivalent.csv, study03_mediator.csv.
Omitted confounder

Drop one common cause and 0.173 becomes 0.495

The true path is 0.20. With the confounder in the model it is estimated at 0.173; drop the confounder and the same path reads 0.495, a bias of +0.295. The trimmed model has CFI 1.000 and RMSEA 0.000, so it fits better than the correct one.

Direction and equivalence

Reverse the arrow and the fit is CFI 0.9999

Data generated as injuries driving climate, fitted the other way round, returns -0.570 with p below 0.001 and RMSEA 0.004. Cause to effect and effect to cause give identical output: beta 0.4927, SE 0.0159, chi-square 7.696 on 8 df.

What this means

All four wrong models clear the conventional cutoffs

Omitting a mediator moves the same coefficient from the direct effect 0.182 to the total effect 0.442, and the fit still passes. A fit index only says the implied covariance matrix is close to the observed one, and several causal structures imply the same matrix. A good fit is not evidence that the structure is right.

With rare events, 680 alerts for every real one

At a base rate of 9.2 events per 100,000 worker-shifts, the metric on the scoreboard and the experience of the supervisor stop agreeing.

Scene 05

With rare events, 680 alerts for every real one

Source: results/study02_alert_burden.csv, study02_discrimination.csv, study02_imbalance.csv, study02_sizing.csv.
Burden

680 alerts for every one real event

That is a model with 80 percent sensitivity and 95 percent specificity at a recordable rate of 9.2 events per 100,000 worker-shifts. Push specificity to 99.99 percent and the burden falls to 2.1 alerts per event.

Base rate

AUC moves 0.007 while alerts per event move 39-fold

The same generating signal gives AUC 0.685 at a 30 percent base rate and 0.692 at 0.5 percent. Precision in the top five percent falls from 0.629 to 0.016 and alerts per true event climb from 1.6 to 62. AUC cannot see the base rate your sites actually have.

Rebalancing and sizing

Undersampling overstates risk by a factor of 88

Trained as collected, AUC is 0.711 and predicted risk matches observed to within two percent. Undersample to a 50 percent event rate and AUC is 0.707, unchanged, while the Brier score goes from 0.005 to 0.230. Sizing is the same story: ten percent relative precision needs 384 events, or 4,175,499 worker-shifts. Ask any alerting vendor for alerts per true event, not AUC.

The same fitted coefficients can mean an effect of 0.45 or of 0.281

Three causal structures return the same fitted numbers and three different answers to what happens if you act.

Scene 06

The same fitted coefficients can mean an effect of 0.45 or of 0.281

Source: results/study04_ambiguity.csv.
Three worlds

Three worlds return the same four coefficients

The fitted vector is [0.448, 0.306, -0.254, -0.200] in world one, [0.453, 0.307, -0.260, -0.208] in world two and [0.448, 0.303, -0.246, -0.203] in world three. Rounded for a slide, all three are the same slide.

Different answers

The effect of acting is 0.45 in world one and 0.281 in world three

Only in world one is the coefficient on unsafe acts the effect of intervening on unsafe acts. In world two the reported -0.260 on system condition is a direct effect and the total effect is -0.070. In world three an unmeasured common cause sits behind unsafe acts and risk, so intervening returns 0.281 while the regression still prints 0.448.

Scale and instrument

4.6 percent of units cross the audit flag when the scale changes

Crew A scores 2.55 and crew B 3.15 on the same raw inputs at both sites, yet standardizing within site puts crew A ahead at site 1 and crew B ahead at site 2, because overtime varies more at one of them. Fit the measurement model by site and one loading drops from 0.769 to 0.301 and composite reliability from 0.823 to 0.675. Forcing site A's model onto site B sends audits to a different 4.6 percent of units.

Limits

What this does not establish

Data

Every dataset on this page is synthetic, generated from a known model by simulations/dgp.py. That is the design, not a shortcut: the failures shown here are only visible when the true answer is known. Nothing here describes a real site, a real workforce or a real rate.

Scope

These are simulation studies of estimator behaviour. They do not show that any particular published safety model is wrong. They show which claims a model of this shape can and cannot support at a given sample size.

Precision

The recovery study uses 300 replications per cell; the bootstrap study uses 40, with a Monte Carlo standard error on coverage of about 0.034. Read the bootstrap coverage figures as approximately 0.95, not exactly.

Status

No manuscript from this work has been submitted or peer reviewed. The code, the generating models and the result tables are all in the repository so the numbers can be regenerated and disagreed with.

Monday

One thing to do on Monday

Take the last risk model your team presented. Ask for the sample size and the interval method behind the coefficient that drove the decision. If the interval came from the analytic standard error, ask for a bootstrap interval instead, and ask what would have to be true about the causal structure for that coefficient to be the effect of acting on it.

  • 1 Get the n. Compare it against 369 for a usable path and 6884 for a claim that one path beats another.
  • 2 Replace the analytic interval with a bootstrap interval. Expect it to be about 40 percent wider.
  • 3 Write down the one causal structure under which the coefficient answers the question being asked. If you cannot, report it as an association.
  • 4 For any alerting model, ask for alerts per true event at the operating point, not AUC.
Next in the argument · 5 of 7

The capital case

What a safety programme is worth when the benefit is an event that did not happen, and how the discount rate decides the answer.