← Back to the post
Interactive data dictionary

Tutoring and GPA in 35 high schools (DiD tutorial)

A simulated school panel in which 10 of 35 high schools adopt an after-school tutoring program at the same time: the data behind every number in the Python Difference-in-Differences tutorial, as a 2×2 file and an 8-period event-study file.

2
datasets
35
schools
10
treated schools
25.32
DiD estimate (GPA points)

Downloads

Each dataset is available as a labeled Stata .dta and its source file.

⇩ Download all data (ZIP)stata_codebook.do

DatasetGrainRowsStataSource
tutoring_didschool x period (balanced panel)70 × 7tutoring_did.dtatutoring_did.csv
tutoring_dideventschool x period (balanced panel)280 × 8tutoring_didevent.dtatutoring_didevent.csv

Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.

Load directly in code

Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.

Stata

* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
use "${BASE}tutoring_did.dta", clear
describe
notes

Python

!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
df = pd.read_stata(BASE + "tutoring_did.dta")

# load every dataset at once
files = ["tutoring_did", "tutoring_didevent"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}

# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "tutoring_did.dta", "tutoring_did.dta")
df, meta = pyreadstat.read_dta("tutoring_did.dta")

Copy and paste this snippet in Google Colab app. https://colab.research.google.com/notebooks/empty.ipynb

R

# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
df <- read_dta(paste0(BASE, "tutoring_did.dta"))

Overview & sources

Companion data for the tutorial Introduction to Difference-in-Differences (DiD) in Python, which uses the simulated case study of Corral and Yang (2024). In one region, 10 of 35 high schools adopt an after-school tutoring program for low-income students at the same time; the other 25 schools never adopt it. The outcome gpa is the school's average GPA of low-income students. In the 2×2 file, treated schools rise from 60.17 to 96.37 (a naive change of 36.20 points) while comparison schools rise from 71.22 to 82.10 (10.89 points), so the DiD estimate of the average treatment effect on the treated (ATT) is 25.315 GPA points (25.32 from the rounded means). The event-study file follows the same 35 schools for 8 periods with adoption in period 5; its event-time coefficients are 0.34, −0.32 and 0.59 before adoption and 25.03 to 25.70 after.

Two files, two separate simulations. tutoring_did is the 2×2 design: 35 schools × 2 periods = 70 rows, keyed by id × time (time 1 = before, 2 = after the program). tutoring_didevent is the event-study panel: the same 35 school ids × 8 periods = 280 rows, keyed by id × time, with treatment from period 5 on. Both panels are strongly balanced and the treated schools are id 26–35 in both. The 2×2 file is not two periods of the event-study file: its GPA values do not appear in it, which is why collapsing the event study to a 2×2 (Exercise 2 in the post) gives 24.90 rather than 25.32.

Data sources

SourceProvidesReference / URL
Corral and Yang (2024)The simulated tutoring case study: 35 high schools, 10 adopting an after-school tutoring program, GPA of low-income students, the 2×2 and event-study designsCorral, D., & Yang, M. (2024). An introduction to the difference-in-differences design in education policy research. Asia Pacific Education Review, 25(3), 663–672. https://doi.org/10.1007/s12564-024-09959-0
QuaRCS-lab data-openThe original Stata files tutoring_did.dta and tutoring_didevent.dta that the post loadshttps://github.com/quarcs-lab/data-open/tree/master/isds
This post (CSV copies)tutoring_did.csv and tutoring_didevent.csv: the same values as the original .dta files, written at full precision so Python, R and Stata read identical numbers (used by the cheat sheets and analysis.do)Mendez, C. (2026). https://carlos-mendez.org/post/python_did101/

Cite this data

Please cite this dataset as follows.

APA

Mendez, C. (2026). Tutoring and GPA in 35 high schools (DiD tutorial) [Data set]. https://carlos-mendez.org/post/python_did101/

Corral, D., & Yang, M. (2024). An introduction to the difference-in-differences design in education policy research. Asia Pacific Education Review, 25(3), 663–672. https://doi.org/10.1007/s12564-024-09959-0

BibTeX

@misc{mendez2026pythondid101,
  author       = {Mendez, Carlos},
  title        = {Tutoring and GPA in 35 high schools (DiD tutorial)},
  year         = {2026},
  howpublished = {\url{https://carlos-mendez.org/post/python_did101/}},
  note         = {Data set}
}

@article{corral2024did,
  author  = {Corral, Daniel and Yang, Minseok},
  title   = {An introduction to the difference-in-differences design in education policy research},
  journal = {Asia Pacific Education Review},
  volume  = {25}, number = {3}, pages = {663--672}, year = {2024},
  doi     = {10.1007/s12564-024-09959-0}
}

Variable explorer search & filter all 8 variables

Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.

VariableTypeDistributionLabelDefinitionUnitsIn filesSource
female_share#continuousmin 0.471 | median 0.527 | max 0.57Share of female students (covariate)Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate.share (0-1)tutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
gpa#continuousmin 59.4 | median 76.3 | max 99.2Average GPA of low-income students (outcome)The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file).GPA pointstutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
id#identifier–School identifier (1-35)High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program.idtutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
post#dummyshare coded 1 = 0.500Post-adoption period (1 = yes)1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file.0/1tutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
time#continuousmin 1 | median 1.5 | max 2Period (2x2: 1-2; event study: 1-8)Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5.periodtutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
timeToTreat#continuousmin -4 | median -0.5 | max 3Periods relative to adoption (treated schools only)Event time for treated schools: time - 5, from -4 to 3 (0 = first treated period, -1 = the reference period in the post). Missing for comparison schools, which are never treated.periodstutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
treated#dummyshare coded 1 = 0.286Treatment group: school adopts tutoring (1 = yes)1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school.0/1tutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open
txp#dummyshare coded 1 = 0.143Treated x post: school is under the program (1 = yes)The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate.0/1tutoring_did, tutoring_dideventCorral and Yang (2024), via quarcs-lab/data-open

Cross-file variable index

Which file each variable appears in (● = present).

Variabletutoring_didtutoring_didevent
female_share●●
gpa●●
id●●
post●●
time●●
timeToTreat●
treated●●
txp●●

Construction & formulas

Design variables

Estimates the post computes from these files

The datasets

Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.

expand to search (Ctrl/⌘+F) or print across all datasets

school x period (balanced panel)  70 × 7 · 2 periods (before, after) · 35 high schools (10 treated, 25 comparison) x 2 periods = 70 rows

Panel key: id x time · Input for the post's 2x2 analysis: naive before-after comparison, manual double difference, OLS interaction, two-way fixed effects, the covariate check and the four standard-error types.

Variable dictionary

VariableLabelDefinitionConstructionUnitsSourceCoverage
id identifierSchool identifier (1-35)High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program.As provided in the original .dta files.idCorral and Yang (2024), via quarcs-lab/data-open35 schools
time continuousPeriod (2x2: 1-2; event study: 1-8)Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5.As provided in the original .dta files.periodCorral and Yang (2024), via quarcs-lab/data-openall rows
treated dummyTreatment group: school adopts tutoring (1 = yes)1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school.1 if id is 26-35.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
post dummyPost-adoption period (1 = yes)1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file.post = (time == 2) in the 2x2 file; post = (time >= 5) in the event-study file.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
txp dummyTreated x post: school is under the program (1 = yes)The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate.txp = treated x post.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
gpa continuousAverage GPA of low-income students (outcome)The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file).Simulated by Corral and Yang (2024); stored as single-precision float in the original .dta, written at full precision in the CSV.GPA pointsCorral and Yang (2024), via quarcs-lab/data-openall rows
female_share continuousShare of female students (covariate)Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate.Simulated by Corral and Yang (2024); single-precision float in the original .dta.share (0-1)Corral and Yang (2024), via quarcs-lab/data-openall rows

Distribution & statistics (click a header to sort)

VariableDistributionCoverageNDistinctMinMeanMedianMaxSD
id–100%7035—————
timemin 1 | median 1.5 | max 2100%7021.001.501.502.000.504
treatedshare coded 1 = 0.286100%70200.28601.000.455
postshare coded 1 = 0.500100%70200.5000.5001.000.504
txpshare coded 1 = 0.143100%70200.14301.000.352
gpamin 59.4 | median 76.3 | max 99.2100%707059.3977.1276.2799.1510.88
female_sharemin 0.471 | median 0.527 | max 0.57100%70700.4710.5280.5270.5700.027

school x period (balanced panel)  280 × 8 · 8 periods; adoption in period 5 · 35 high schools (10 treated, 25 comparison) x 8 periods = 280 rows

Panel key: id x time · Input for the post's event study: one coefficient per period relative to adoption (t = -4 ... 3, reference t = -1).

Variable dictionary

VariableLabelDefinitionConstructionUnitsSourceCoverage
id identifierSchool identifier (1-35)High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program.As provided in the original .dta files.idCorral and Yang (2024), via quarcs-lab/data-open35 schools
time continuousPeriod (2x2: 1-2; event study: 1-8)Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5.As provided in the original .dta files.periodCorral and Yang (2024), via quarcs-lab/data-openall rows
treated dummyTreatment group: school adopts tutoring (1 = yes)1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school.1 if id is 26-35.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
gpa continuousAverage GPA of low-income students (outcome)The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file).Simulated by Corral and Yang (2024); stored as single-precision float in the original .dta, written at full precision in the CSV.GPA pointsCorral and Yang (2024), via quarcs-lab/data-openall rows
female_share continuousShare of female students (covariate)Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate.Simulated by Corral and Yang (2024); single-precision float in the original .dta.share (0-1)Corral and Yang (2024), via quarcs-lab/data-openall rows
post dummyPost-adoption period (1 = yes)1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file.post = (time == 2) in the 2x2 file; post = (time >= 5) in the event-study file.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
txp dummyTreated x post: school is under the program (1 = yes)The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate.txp = treated x post.0/1Corral and Yang (2024), via quarcs-lab/data-openall rows
timeToTreat continuousPeriods relative to adoption (treated schools only)Event time for treated schools: time - 5, from -4 to 3 (0 = first treated period, -1 = the reference period in the post). Missing for comparison schools, which are never treated.timeToTreat = time - 5 if treated == 1; missing otherwise.periodsCorral and Yang (2024), via quarcs-lab/data-open80 of 280 rows (treated schools)

Distribution & statistics (click a header to sort)

VariableDistributionCoverageNDistinctMinMeanMedianMaxSD
id–100%28035—————
timemin 1 | median 4.5 | max 8100%28081.004.504.508.002.30
treatedshare coded 1 = 0.286100%280200.28601.000.453
gpamin 60.1 | median 78.5 | max 108100%28028060.0880.1478.53107.712.20
female_sharemin 0.47 | median 0.524 | max 0.57100%2802800.4700.5210.5240.5700.028
postshare coded 1 = 0.500100%280200.5000.5001.000.501
txpshare coded 1 = 0.143100%280200.14301.000.351
timeToTreatmin -4 | median -0.5 | max 329%808-4.00-0.500-0.5003.002.31

Known limitations & caveats