Same regression, same data — and shock half-lives of 1.5, 9, or 18 years
How much of this year’s employment shock is still visible next year?
One number answers it: \(\rho\), the coefficient on lagged employment.
Same 140-firm UK panel. Three answers: \(0.626\), \(0.927\), \(0.962\).
One of them is defensible. None of them announces which.
Firms orbit their own levels — a fixed effect and persistence at once
Each blue line is one firm, 1976–1984. Lines are parallel-ish (a firm fixed effect) and smooth (high persistence). The orange median dips after 1980 — the UK recession the year dummies absorb.
The model commits the one sin ordinary panel methods cannot forgive
Last year’s employment depends on \(\alpha_i\) by construction — so the regressor is correlated with the error, and no control can fix it.
Each estimator fails informatively — and hands off to the next
OLS and fixed effects fail in known directions — and bracket the truth
Anderson-Hsiao IV: consistent, and useless
Difference GMM: passes every test, hugs the wrong bound
System GMM: the defensible \(\hat\rho = 0.927\) — then we stress-test it
The Investigation
Act II
Pooled OLS and fixed effects fail in opposite, known directions
Pooled OLS — biased UP
Ignores \(\alpha_i\): it stays in the error
\(n_{i,t-1}\) is positively correlated with \(\alpha_i\)
The lag gets credit for the firm effect’s work
\(\hat\rho = 0.962\) (SE 0.008) — near a unit root
Fixed effects — biased DOWN
Demeaning removes \(\alpha_i\) exactly…
…but the firm mean contains future shocks
Nickell bias, order \(1/T\) — and \(T\) is 7–9 here
\(\hat\rho = 0.626\) (SE 0.052)
Two wrong answers bracket the truth: any consistent estimate must land in [0.626, 0.962]
Bond’s (2002) bracket: OLS marks the ceiling, FE the floor. The shaded band is the playing field for every estimator that follows.
Anderson-Hsiao IV is consistent — and useless: 1.233 with a CI 1.87 wide
First-difference away \(\alpha_i\), then instrument \(\Delta n_{i,t-1}\) with the level \(n_{i,t-2}\):
Estimate
SE
95% CI
\(\hat\rho\) (Anderson-Hsiao 2SLS)
1.233
0.478
[0.296, 2.170]
The CI contains the whole bracket, the unit root, and explosive dynamics — all at once.
Under sequential exogeneity, every lag dated \(t-2\) or earlier is a valid instrument
\[E\big[\,n_{i,t-s}\,\Delta\varepsilon_{it}\,\big] = 0 \qquad \text{for all } s \ge 2\]
Employment two or more years ago carries no information about this year’s change in shocks — dozens of such moment conditions, combined optimally by GMM.
Two command strings run the whole GMM ladder in pydynpd
Even for a persistent series, last year’s change is informative about this year’s level — at the price of one new assumption, mean stationarity.
Ninety-three percent of an employment shock survives into next year
0.927
system GMM, two-step, 32 collapsed instruments (SE 0.079) — inside the bracket, upper half
The headline’s diagnostics are textbook-clean
Two-step system GMM (collapsed)
\(\hat\rho\) (L1.n)
0.927
SE (Windmeijer)
0.079
Instruments
32 (vs 140 firms)
AR(1) p
0.000 — rejects, as required
AR(2) p
0.994 — clean
Hansen p
0.462 — comfortable middle
Every box is ticked — which is exactly why two of these three tests deserve a second look.
The diagnostics decoder: two of the three tests are read backwards
Test
Correct reading
Headline value
AR(1) in differences
Must reject — rejection is mechanical good news
p = 0.000 ✓
AR(2) in differences
Must not reject — this validates the \(t-2\) instruments
p = 0.994
Hansen J
Two-tailed in spirit — p < 0.05 is invalid; p near 1 is an overwhelmed test
p = 0.462 ✓
AR(1) rejecting is the model working, not failing; a Hansen p near 1 is the test overwhelmed, not the instruments vindicated.
The Hansen p-value responds to the instrument count, not just validity
Lag window
Collapsed
Instruments
\(\hat\rho\)
Hansen p
2:3
no
68
0.956
0.035
2:3
yes
17
0.921
0.096
2:5
no
95
0.935
0.186
2:99
no
113
0.930
0.235
2:99
yes
32
0.927
0.462
Same model, six ways — five shown: uncollapsed 2:3 is rejected (p = 0.035) while its collapsed twin passes, driven purely by instrument count.
More instruments is not better — proliferation disarms the test that guards you
Hansen p against instrument count: full matrix (steel) vs collapsed (teal), the p = 0.05 rejection line, and the 140-firm Roodman ceiling.
The toolchain replicates the published benchmark digit for digit
Published vignette
Our run
L1.n
0.2710675
0.2710675
Hansen \(\chi^2\)
32.666
32.666
Instruments
42
42
Exact match under a hard assertion — the NumPy-2 compatibility shim perturbs nothing.
Seven estimators, one parameter — only the workflow identifies the winner
Every estimator with its 95% CI against the shaded OLS–FE bracket. Grey defines the band; blue (Anderson-Hsiao) straddles everything; orange (difference GMM) hugs the floor; teal (system GMM) lands in the upper half.
Difference GMM’s table looks as healthy as system GMM’s — only the bracket tells them apart.
Trust, But Verify
Act III
“Your own CI includes the unit root — why believe 0.927?”
Objection. The headline CI [0.773, 1.081] contains \(\rho = 1\), mean stationarity is untestable, and one common \(\rho\) is imposed on all 140 firms.
Response. All true. So the claim is the point estimate and its lower bound — never “employment is stationary.” It survives four checks: the bracket, a 6-cell grid (0.921–0.956), clean AR(2)/Hansen, and an exact replication. The mean-stationarity price is stated out loud.
Six habits that keep a dynamic-panel coefficient defensible
Run OLS and FE first — record the bracket; they are the measuring stick
Diff GMM hugging the FE bound = weak instruments — Hansen and AR(2) do not clear it
Prefer system GMM when persistence is high — and name the price: mean stationarity
AR(1) must reject; AR(2) must not; read Hansen two-tailed
Collapse instruments and report the count relative to the number of groups
Replicate a published benchmark before trusting novel numbers
No single p-value separates 0.927 from 0.679 — the bracket-plus-diagnostics workflow does.