Same workers, same question — but the answer triples depending on the estimator
A simple regression says joining a union raises wages by 7.5 log points. Compare each worker to themselves across years and the number jumps to 21 log points (about 23%).
One dataset, two stories. Which one do you report?
Six estimators on one panel disagree by a factor of three
Six panel estimators with 95% CIs. POLS / Between / RE cluster near 0.07–0.11; FDFE / FE / CRE cluster near 0.21. The classical Hausman test is annotated.
Where we’re going
The data: a balanced 2-period, 2,199-worker wage panel
Between vs within variation — what each estimator can actually use
Seven estimators, escalating discipline: POLS → Between → FD → FE → TWFE → RE → CRE
Hausman and Mundlak — choosing between fixed and random effects
The Investigation
Act II
The lab: 2,199 workers, two years, a perfectly balanced T = 2 panel
Treatment — union membership; only 16.3% of worker-years unionized
Window — restricted to 2010 and 2012, so \(T = 2\) and the panel is balanced
With \(T = 2\), every worker contributes exactly two rows — the cleanest setting to see that first differences and the within estimator are exactly linked: FD without an intercept is FE, and FD with an intercept is two-way FE.
Each estimator chooses which variation to believe
Cross-sectional camp
POLS — ignores the panel
Between — worker means only
RE — GLS-weighted blend
Answers: “union vs non-union workers?”
Within camp
FD / FDFE — period differences
FE / TWFE — demeaned data
CRE / Mundlak — the bridge
Answers: “same worker, switched status?”
Hausman and the Mundlak term are the formal tests for choosing between the two camps.
Before you look: how much of union’s variance is within workers?
Union status has an overall SD of 0.369 and a mean of 16.3%. What share of its variance comes from workers who change status between 2010 and 2012?
A. Less than 10%
B. About 25%
C. About 50%
94% of union variation is between workers — only 6.1% is within
Between vs within variance shares for the four key variables. Union is 93.9% between; schooling is 100% between (zero within).
Only the workers who switch status identify the within estimators
Log-wage trajectories for 30 sampled workers. Teal lines change union status; orange/blue lines never do.
Pooled OLS — the naive baseline — reports a 7.5-log-point premium
Highly significant (\(t \approx 3.25\)) — and almost certainly biased if ability selects out of union jobs.
Before you look: smaller, the same, or larger than pooled OLS?
Pooled OLS gave 0.0750 (SE 0.0231). Now subtract each worker’s 2010 values from the 2012 values and regress the change in log wage on the change in union status.
A. Smaller than 0.075
B. About the same
C. Larger, with a larger standard error
First-differencing erases \(\alpha_i\) and triples the estimate to 0.211
The worker-specific effect \(\alpha_i\) — ability, schooling, gender — cancels in the subtraction. What is left is identified only by workers who changed union status.
FDFE \(= 0.2113\) (SE 0.079); the SE is \(3.4\times\) larger than POLS — the switcher-only signature.
The within transformation demeans the data — and the slope steepens to 0.21
Within transformation: raw scatter with the shallow POLS slope (left); demeaned scatter with the steeper FE slope through the origin (right).
Before you look: does fixed effects match FD’s 0.2113?
With only two periods, demeaning and differencing use the same 73 switchers. Will the within (FE) estimate equal the FD estimate of 0.2113?
A. Exactly equal
B. Slightly different
C. Very different
Three recipes, one number: FD, demeaning, and dummy FE all give 0.2103
fit_fe = pf.feols("lwage ~ union | ID", data=df, vcov="HC1") # absorbed FEfit_dvfe = pf.feols("lwage ~ union + C(ID_str)", data=df, vcov="HC1") # 2,198 dummies# Both → 0.2103; FDFE → 0.2113 (the +0.001 is an intercept-driven year trend)
Within transformation, first-differences, and dummy-variable FE are three recipes for the same dish. Absorption (| ID) is just the fast one.
Before you look: where does two-way FE land?
FE gave 0.2103 and FD gave 0.2113. Adding year effects to FE allows a common wage trend.
A. Exactly equal to FD, 0.2113
B. Equal to FE, 0.2103
C. Somewhere else
Two-way FE absorbs year shocks and lands at 0.2113 — exactly first differences
fit_twfe = pf.feols("lwage ~ union | ID + year", data=df, vcov={"CRV1": "ID"})# Union coefficient: 0.2113 (SE 0.0792)
Absorbing the year effect removes the aggregate wage trend FD’s intercept was capturing. Time-invariant regressors (schooling, female) are silently absorbed.
Random effects bets on no-correlation — and is pulled toward POLS at 0.109
RE is a variance-weighted average of between and within. With only 6% within, it leans toward the between picture — and SE is \(2.7\times\) tighter than FE.
Before you look: what do robust SEs do to the Hausman statistic?
The textbook test uses classical SEs: 0.0509 for FE and 0.0278 for RE. The robust SEs are larger, 0.0812 and 0.0299. Plug the robust ones into the same formula, and H will…
A. Rise, so the test rejects more strongly
B. Fall enough to flip the verdict
C. Fall, but still reject at 5%
The textbook Hausman test rejects RE, if the errors are classical
With classical variances (SE 0.0509 for FE, 0.0278 for RE): \(H = 5.62\), \(p = 0.018\), so we reject RE at 5%.
Plugging the robust SEs into the same formula gives \(H = 1.79\), \(p = 0.180\), but that is not a valid test: \(V_{\mathrm{FE}} - V_{\mathrm{RE}}\) is the variance of the difference only when RE is efficient.
Before you look: what happens to the union coefficient in CRE?
Random effects gave 0.1092. Now add each worker’s mean union status, union_bar, as a regressor and re-estimate random effects.
A. It equals the FE value, 0.2103
B. It stays near 0.11
C. It moves between 0.11 and 0.21
Mundlak recovers the FE coefficient and flags negative selection
Add each worker’s mean union exposure \(\bar{x}_i\), then run RE (random worker effect \(a_i\)). The within coefficient \(\beta\) equals FE; the mean coefficient \(\gamma\) tests whether \(\alpha_i\) correlates with \(x\).
CRE within \(= 0.2103\) (matches FE exactly); Mundlak term \(\gamma = -0.144\), \(p = 0.072\) — borderline, hinting at negative selection.
The Resolution
Act III
Within-worker, joining a union pays 0.21 log points — nearly triple the naive 0.075
0.210
\(\hat\beta_{\mathrm{FE}}\) on union (SE 0.081) — vs pooled OLS 0.075; FDFE, TWFE, and CRE all agree near 0.21
Two camps, three-fold apart — and the gap is selection, not noise
Method
Coef
SE
Variation used
POLS
0.0750
0.0231
all (ignores panel)
Between
0.0662
0.0311
cross-sectional means
RE
0.1092
0.0299
GLS between + within
FDFE
0.2113
0.0792
within differences
FE
0.2103
0.0812
within demeaned
TWFE
0.2113
0.0792
within, net of year effects
CRE
0.2103
0.0703
RE + Mundlak (= within)
Cross-sectional 7–11 log points · within ~21. Standard errors swing inversely — the within camp is noisier but causally cleaner.
Before you look: what happens to the age coefficient under two-way FE?
With controls, each extra year of age raises log wages by about 0.02 in pooled OLS. Two-way FE absorbs both worker and year effects. Under two-way FE, the age coefficient will…
A. Stay near +0.02
B. Change sign
C. Shrink toward zero
Adding controls leaves the two-camp gap intact
Extended models — union, age, schooling, female across POLS / TWFE / RE / CRE. The within premium survives controls.
Does FE make this causal? No — strict exogeneity still carries the weight
Objection. Within estimators just net out fixed traits — they can’t manufacture identification.
Response. Correct. FE/FDFE/TWFE/CRE target the average effect among union switchers only — and only under strict exogeneity given the worker fixed effect.
The 0.21 rests on 73 switchers and on two waves
Joiners vs leavers — two-way FE on stayers + joiners gives 0.345; on stayers + leavers, 0.081
All five waves (2010–2018) — two-way FE falls to 0.040 (SE 0.026), FD with year effects to 0.057
More waves, more within variation — the within share of union rises from 6.1% to 16.1%, yet the premium shrinks
The within estimate is credible about what it identifies, not about how general it is: 0.21 is a 2010–2012 number for a small, possibly unusual group.
With robust errors, lead with CRE/Mundlak and its built-in test
p = 0.072
Robust Mundlak term (0.106 clustered): borderline. The textbook Hausman rejects (p = 0.018) only under classical errors.
Let the within variation, not the pooled average, tell you what a treatment does.