What does panel data buy you?
Pooled OLS on a worker wage panel says joining a union raises wages by about 7.5 log points (roughly 7.5%). The within (fixed-effects) estimator on the very same data says about 21 log points (about 23%) — almost three times larger. That factor-of-three gap is the empirical signature of selection on unobservables, and panel methods exist to peel it apart.
This app turns the post's three central ideas into knobs you can move. Watch the within transformation demean three workers in real time. Sweep panel sample size and unit heterogeneity to see how pooled OLS drifts from the truth while fixed effects stays close. Toggle the seven canonical estimators in the forest plot. And drive the Hausman statistic yourself by changing the FE/RE coefficient gap.
The within transformation, animated
Three workers, two periods each. Alice (steel) never joins the union. Bob (teal) is a low-wage worker who joins the union between periods. Carla (orange) is always in the union. The dashed grey line is the pooled OLS fit — shallow, because Bob's low wages pull the union and non-union averages together, just as negative selection does in the post's data. Wait a few seconds: the points slide as each worker's two-year mean is subtracted. After demeaning, only Bob moves — and the FE slope (orange) is much steeper.
The animation cycles between raw and demeaned coordinates on a 10-second loop. Demeaning removes two of three workers from the slope (Alice and Carla collapse onto union = 0) because they never switched. Only switchers identify the FE coefficient — that is the post's central asymmetry.
Between vs within: how much variation does FE have to work with?
For each variable in the post's wage panel (N = 2,199 workers; T = 2), we decompose total variance into a between part (variation across workers) and a within part (variation over time inside a worker). FE only uses the within slice.
Schooling has zero within-variation in this two-period sample — no worker's education changes between 2010 and 2012, so FE mechanically drops schooling. Union has only 6% within-variation: that thin slice is what FE uses to identify the union premium. Less data, different (and arguably cleaner) parameter.
Panel DGP Simulator
Simulate a panel with unit fixed effects. Slide unit heterogeneity up and watch POLS drift away from the truth while FE stays put. Run 100 simulations for a bias-variance picture.
Seven-Method Forest Plot
The post's headline figure, interactively. Toggle outcomes and methods. POLS/Between/RE land at 7-11 log points; FDFE/FE/TWFE/CRE land at 21. Three-fold gap, one dataset.
Hausman vs Mundlak
Drive the Hausman chi-square yourself by changing the FE/RE coefficient gap and the standard errors. See why the textbook test rejects RE, why plugging in robust SEs flips the verdict, and why the Mundlak test is the robust check.
Glossary (open a card if a term is unfamiliar)
Within transformation
Between estimator
Fixed effects (FE)
First differences (FDFE)
Two-way FE (TWFE)
Random effects (RE)
Mundlak / CRE
Hausman test
Panel DGP Simulator — POLS vs FE on simulated data
Simulate a panel with N workers observed over T periods. Each worker has their own intercept (the unit fixed effect alpha_i) and a treatment x_it that is correlated with that intercept — so POLS will be biased. The true treatment coefficient is beta = 0.21 (matching the post's FE estimate). Slide unit heterogeneity up and watch POLS drift away from 0.21 while FE stays close.
Pooled OLS
Ignores panel structure. Each row is treated as independent.
Fixed Effects
Unit demeaning. Within-only variation.
What to look for
- Slide unit heterogeneity to 2.0. POLS drifts far above the truth because high-FE workers cluster at high x. This is omitted-variable bias made visible; in this simulator the bias is upward, whereas in the post's data it is downward (POLS 0.075 below FE 0.21), because there low-wage workers select into unions.
- Slide selection ρ to 0. POLS becomes unbiased — the two methods agree. This is the RE assumption: no correlation between alpha_i and x.
- Slide T to 8. More periods give FE more within variation to work with, so the spread of the FE estimates shrinks. With T = 2 (the post's choice), FE is consistent but high-variance.
- True coefficient 0.21 is the dashed grey line. POLS scatters above the truth; FE clusters near it.
Bias vs variance across 100 simulations
One draw is noisy. Run the panel simulation 100 times with the current sliders to see whether POLS bias is systematic and how variance compares to FE.
The post's forest plot, interactively
These numbers come straight from basic_models_comparison.csv
and extended_models_comparison.csv in the post's folder.
Seven methods on the basic model (union as the lone regressor) and four
methods with controls (age, schooling, female, year effects; TWFE
absorbs schooling and female). Toggle
outcomes and methods to compare. Hover any point for SE, 95% CI, and
sample size.
What to look for
- Two clear camps on "union (basic)". POLS/Between/RE land at 7–11 log points; FDFE/FE/TWFE/CRE land at 21. That factor-of-three gap is the empirical signature of selection on unobservables.
- CRE equals FE exactly (0.2103 to four decimals). Mundlak's algebra is exact: add the unit means and RE recovers the within coefficient on the original variable.
- Schooling is absorbed under TWFE. No row appears for TWFE on schooling because schooling has zero within-variation. The other three methods recover 0.111.
- The age coefficient flips sign in the within models (+0.021 in POLS and RE, −0.058 in TWFE and CRE; with year effects CRE reproduces TWFE exactly). Age is not exactly collinear with the year dummy (it rose by 2 years for 1,885 workers, by 1 for 164, and by 3 for 150), so the TWFE age coefficient is identified only from the 314 workers whose age did not rise by exactly two years — a fragile estimate, not a real age–wage profile.
Outcomes
Methods
Why does the union premium triple?
POLS asks: "how do union and non-union workers compare?" Their answer (7.5 log points) is biased if union members are systematically different from non-members on unobservables (ability, motivation, industry).
FE asks a sharper question: "what happens when the same worker switches union status?" Only the workers who actually changed union status between 2010 and 2012 contribute. The answer (21 log points) is what we would report if we trusted that nothing else changes for a worker over those two years.
The triple is consistent with negative selection into unions: higher-ability workers are less likely to be in unions in this sample, so cross-sectional comparisons understate the within-worker payoff to joining a union.
But the 0.21 is local and fragile. It rests on the 73 workers (3.3%) who switched. Joiners alone give 0.345 and leavers alone 0.081, so the pooled 0.21 averages two very different responses. With all five waves (2010–2018), two-way FE falls to 0.040 (SE 0.026) and FD with year effects to 0.057: the 0.21 is a 2010–2012 number.
Hausman vs Mundlak — when should you trust RE?
The Hausman statistic compares the FE and RE coefficient estimates,
weighted by the difference in their variances:
H = (beta_FE − beta_RE)² / (V_FE − V_RE).
Large H ⇒ FE and RE disagree ⇒ reject RE. The catch: the
formula is valid only with classical (non-robust) variances,
because V_FE − V_RE is the variance of the difference only when RE is
efficient. With robust errors, use the Mundlak test instead.
What to look for
- Snap to post values (β̂_FE = 0.2103, β̂_RE = 0.1092, classical SE_FE = 0.0509, SE_RE = 0.0278). H = 5.62, p = 0.018: the textbook test rejects RE.
- Plug in the robust SEs (SE_FE = 0.0812, SE_RE = 0.0299; the sliders land on 0.081 and 0.030). H falls to about 1.8 and p rises to about 0.18, so the verdict flips. That number is not a valid test, because with robust errors V_FE − V_RE is no longer the variance of the difference; it does show how a larger SE_FE drains the test of power.
- Set β̂_FE = β̂_RE. H drops to 0 and p = 1. When the two estimators agree, there is no evidence against RE — which is exactly the test's logic.
- The Mundlak alternative (post §14) is valid with robust errors: p = 0.072 with White SEs and p = 0.106 clustered by worker. Borderline, not a rejection at 5%.
Why lead with Mundlak when the errors are not classical?
Mundlak's specification adds x_bar (the unit mean of the
time-varying regressor) to an RE regression. The coefficient on
x_bar is a direct test of correlation between the unit
effect and the regressor — the very assumption RE relies on. Under
classical errors it is equivalent to the Hausman test; its advantage
is that its standard error can be made robust to heteroskedasticity
and clustering, which the Hausman formula cannot.
In the post's data, the textbook Hausman test gives p = 0.018 (reject),
but it assumes classical errors. The robust Mundlak test on
union_bar gives p = 0.072 (0.106 clustered): borderline.
In this balanced panel the Mundlak coefficient, −0.144, is exactly the
between estimate minus the FE estimate.
For a practitioner, CRE/Mundlak is usually the right
specification to lead with: you get the FE coefficient on
time-varying treatments, the RE structure for keeping schooling and
gender, and a built-in spec test that stays valid with robust errors.
Connecting back to Tab 2
The Tab 2 simulator lets you drive the FE-vs-POLS bias through the selection ρ slider. The Hausman test on that simulated data would reject RE if ρ is large enough, and more easily when SE_FE is small. The simulator does not compute the test itself; the explorer above shows how the verdict depends on SE_FE for a fixed coefficient gap.