Introduction to Difference-in-Differences in Python

Does after-school tutoring raise GPA? Separating the program from the trend

25.32ATT · GPA points (DiD)
+43%naive overstates the DiD effect
0significant pre-trends

Carlos Mendez

Nagoya University (GSID)

October 3, 2026

The Tension

Act I

Tutored schools’ GPA jumped 36 points — case closed?

A district launched after-school tutoring in 10 of its 35 high schools. One year later, average GPA in those schools rose from 60.17 to 96.37.

A 36.20-point leap. But the 25 untreated schools also improved — from 71.22 to 82.10. How much of the jump is really the program?

The naive before-after credits the entire 36-point rise to the program

Treated-group GPA rising from 60.17 to 96.37: the naive estimator credits the full 36.20-point increase to tutoring, ignoring everything else that changed.

The plan: from a naive before-after to an 8-period event study

  • The lab: 35 schools in 2 periods, a clean 2×2 design
  • Why the naive estimate overstates, and how a comparison group fixes it
  • Manual double-differencing, then OLS and two-way fixed effects in PyFixest
  • Four standard errors: does inference change the verdict?
  • An 8-period event study, then five tempting misreadings

The Investigation

Act II

A clean 2×2 lab: 35 schools, 2 periods, simultaneous adoption

Panel structure: 10 treated schools (orange) switch on after the intervention; 25 comparison schools (steel blue) never do. Timing is simultaneous and treatment, once on, stays on.

Before you look: what happened in the schools without tutoring?

Tutored schools rose 36.20 points. What did the 25 comparison schools do, and where does that leave DiD?

A.  They stayed flat, so DiD is about 36

B.  They rose by about 11 points, so DiD is about 25

C.  They fell, so DiD is larger than 36

The comparison group reveals an 11-point trend the program never caused

Group Pre Post Change
Comparison (25) 71.22 82.10 +10.88
Treated (10) 60.17 96.37 +36.20

The comparison group’s +10.88 is the secular trend: the rise that would have happened anyway.

DiD is a double difference: 36.20 minus 10.88 equals 25.32

\[DiD = \Big(\bar{Y}^{post}_{treat} - \bar{Y}^{pre}_{treat}\Big) - \Big(\bar{Y}^{post}_{ctrl} - \bar{Y}^{pre}_{ctrl}\Big)\]

\[DiD = (96.37 - 60.17) - (82.10 - 71.22) = 36.20 - 10.88 = 25.32\]

Subtract the trend the comparison group reveals; what is left is the program’s effect.

The counterfactual: where treated schools would have landed without tutoring

Three lines: comparison (steel, 71.22→82.10), treated (orange, 60.17→96.37) and the counterfactual (teal dashed, 60.17→71.05). The 25.32 gap between the actual and counterfactual treated outcome is the ATT.

A treated×post interaction recovers the same 25.32, and every coefficient is a group mean

fit_ols = pf.feols("gpa ~ treated + post + txp", data=df, vcov="HC1")
print(fit_ols.summary())
Coefficient Estimate Maps to
Intercept 71.215 comparison pre-mean
treated −11.049 baseline level gap
post +10.886 common time trend
txp 25.315 the DiD estimate

Before you look: replace treated and post with fixed effects

Swap them for school and period fixed effects, and cluster by school. What changes?

A.  The coefficient stays at 25.315; only the standard error changes

B.  Both the coefficient and the standard error change

C.  Nothing changes: same coefficient, same standard error

Two-way fixed effects: the same 25.315, now with school-clustered errors

\[Y_{it} = \beta\,(\text{Treat}_i \times \text{Post}_t) + \gamma_i + \vartheta_t + \varepsilon_{it}\]

fit_twfe = pf.feols("gpa ~ txp | id + time", data=df, vcov={"CRV1": "id"})

Coefficient unchanged at 25.315; the SE moves from 0.615 (HC1) to 0.585 (clustered by school).

Before you look: does female_share move the estimate?

The share of female students varies a little within schools. Add it to the model: what happens to 25.315?

A.  It falls by more than one GPA point

B.  It rises by more than one GPA point

C.  It barely moves: less than 0.1 points

Three specifications, one answer: 25.315 to 25.328

The txp estimate across OLS/HC1, TWFE/CRV1 and TWFE + covariate/CRV1: all clustered tightly around 25.32, with narrow intervals far from zero.

Stability is the reassuring part; an insignificant covariate does not prove the fixed effects capture everything.

Before you look: which standard error is largest?

Same model, same coefficient (25.315), four variance estimators: iid, HC1, CRV1 (clustered by school) and CRV3 (a leave-one-school-out jackknife). Which reports the largest standard error?

A.  iid

B.  HC1

C.  CRV3

Inference choices barely move the needle when the signal is this strong

Standard errors across iid, HC1, CRV1 and CRV3. CRV3 is largest (0.637); HC1 and CRV1 are nearly identical (0.585). All four give overwhelmingly significant results.

Here the design, not the variance estimator, decides the verdict: every t-statistic is above 39.

The Resolution

Act III

The program raises GPA by 25.32 points, not 36.20

25.32

ATT, stable across specifications (25.315–25.328) · the naive 36.20 overstated it by 43%

The event-study panel: treatment switches on in period 5

Panel structure for the event study: 35 schools across 8 periods. Treatment begins in period 5; the 10 treated schools switch from light to dark orange, and the 25 comparison schools stay steel blue.

Before you look: what should the event study show?

Eight periods, one coefficient per period, each relative to t = −1. If parallel trends holds, what do we see?

A.  Leads near zero; lags jump to about 25 and stay roughly flat

B.  Leads near zero; lags climb steadily from about 10 to about 40

C.  Leads already around 25

Leads near zero and an effect from the first treated period: a textbook event study

Event-study coefficients by period relative to treatment. Leads (t = −4 to −2) sit near zero with intervals covering zero; lags (t = 0 to 3) jump to about 25 and stay flat.

The numbers behind the picture: quiet leads, loud lags

Period Estimate 95% CI Significant?
t = −4 0.34 [−0.47, 1.16] no
t = −3 −0.32 [−1.22, 0.57] no
t = −2 0.59 [−0.27, 1.45] no
t = 0 25.03 [24.12, 25.93] yes
t = 3 25.70 [24.08, 27.32] yes

Three near-zero leads are consistent with parallel trends; four large lags trace a stable effect.

Five tempting misreadings

Common misconceptions

“25.32 is the program’s effect for every school”: it is the ATT

Myth. The estimate tells us what tutoring would do in any school.

Truth. DiD identifies the average effect on the 10 treated schools. It says nothing about how the 25 comparison schools would have responded, which could differ if schools chose the program because they expected to benefit.

“Adding controls always makes DiD more credible”: design does, controls may not

Myth. The more covariates, the more believable the estimate.

Truth. Adding female_share moved the estimate by 0.013 (25.315 to 25.328): the credibility comes from the design. A covariate the program itself can change is a bad control and can bias the estimate.

“The SE choice never matters”: it can decide the result for smaller effects

Myth. All four standard errors agreed here, so the choice of estimator is irrelevant.

Truth. They agree because the effect is about 40 standard errors from zero. An estimate of 1.19 points would pass with CRV1 but fail with CRV3 (threshold 1.30): with small effects, the choice decides the verdict.

Does precision make this causal? No — the assumptions still carry it

Objection. The estimate is precise and the pre-trends are flat, so DiD has proven that tutoring caused the gain.

Response. Precision is not identification. The 25.32 ATT rests on parallel trends and SUTVA; flat pre-trends support them but cannot prove them. The data are simulated: in the field, expect imperfect pre-trends, smaller effects and R² well below 0.99.

A credible comparison group, not the variance estimator, is what makes the effect causal.