Downloads
Each dataset is available as a labeled Stata .dta and its source file.
⇩ Download all data (ZIP)stata_codebook.do
| Dataset | Grain | Rows | Stata | Source |
|---|---|---|---|---|
tutoring_did | school x period (balanced panel) | 70 × 7 | tutoring_did.dta | tutoring_did.csv |
tutoring_didevent | school x period (balanced panel) | 280 × 8 | tutoring_didevent.dta | tutoring_didevent.csv |
Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.
Load directly in code
Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.
Stata
* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
use "${BASE}tutoring_did.dta", clear
describe
notesPython
!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
df = pd.read_stata(BASE + "tutoring_did.dta")
# load every dataset at once
files = ["tutoring_did", "tutoring_didevent"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}
# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "tutoring_did.dta", "tutoring_did.dta")
df, meta = pyreadstat.read_dta("tutoring_did.dta")Copy and paste this snippet in Google Colab app. https://colab.research.google.com/notebooks/empty.ipynb
R
# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_did101/data/"
df <- read_dta(paste0(BASE, "tutoring_did.dta"))Overview & sources
Companion data for the tutorial Introduction to Difference-in-Differences (DiD) in Python, which uses the simulated case study of Corral and Yang (2024). In one region, 10 of 35 high schools adopt an after-school tutoring program for low-income students at the same time; the other 25 schools never adopt it. The outcome gpa is the school's average GPA of low-income students. In the 2×2 file, treated schools rise from 60.17 to 96.37 (a naive change of 36.20 points) while comparison schools rise from 71.22 to 82.10 (10.89 points), so the DiD estimate of the average treatment effect on the treated (ATT) is 25.315 GPA points (25.32 from the rounded means). The event-study file follows the same 35 schools for 8 periods with adoption in period 5; its event-time coefficients are 0.34, −0.32 and 0.59 before adoption and 25.03 to 25.70 after.
tutoring_did is the 2×2 design: 35 schools × 2 periods = 70 rows, keyed by id × time (time 1 = before, 2 = after the program). tutoring_didevent is the event-study panel: the same 35 school ids × 8 periods = 280 rows, keyed by id × time, with treatment from period 5 on. Both panels are strongly balanced and the treated schools are id 26–35 in both. The 2×2 file is not two periods of the event-study file: its GPA values do not appear in it, which is why collapsing the event study to a 2×2 (Exercise 2 in the post) gives 24.90 rather than 25.32.
Data sources
| Source | Provides | Reference / URL |
|---|---|---|
| Corral and Yang (2024) | The simulated tutoring case study: 35 high schools, 10 adopting an after-school tutoring program, GPA of low-income students, the 2×2 and event-study designs | Corral, D., & Yang, M. (2024). An introduction to the difference-in-differences design in education policy research. Asia Pacific Education Review, 25(3), 663–672. https://doi.org/10.1007/s12564-024-09959-0 |
| QuaRCS-lab data-open | The original Stata files tutoring_did.dta and tutoring_didevent.dta that the post loads | https://github.com/quarcs-lab/data-open/tree/master/isds |
| This post (CSV copies) | tutoring_did.csv and tutoring_didevent.csv: the same values as the original .dta files, written at full precision so Python, R and Stata read identical numbers (used by the cheat sheets and analysis.do) | Mendez, C. (2026). https://carlos-mendez.org/post/python_did101/ |
Cite this data
Please cite this dataset as follows.
APA
Mendez, C. (2026). Tutoring and GPA in 35 high schools (DiD tutorial) [Data set]. https://carlos-mendez.org/post/python_did101/
Corral, D., & Yang, M. (2024). An introduction to the difference-in-differences design in education policy research. Asia Pacific Education Review, 25(3), 663–672. https://doi.org/10.1007/s12564-024-09959-0BibTeX
@misc{mendez2026pythondid101,
author = {Mendez, Carlos},
title = {Tutoring and GPA in 35 high schools (DiD tutorial)},
year = {2026},
howpublished = {\url{https://carlos-mendez.org/post/python_did101/}},
note = {Data set}
}
@article{corral2024did,
author = {Corral, Daniel and Yang, Minseok},
title = {An introduction to the difference-in-differences design in education policy research},
journal = {Asia Pacific Education Review},
volume = {25}, number = {3}, pages = {663--672}, year = {2024},
doi = {10.1007/s12564-024-09959-0}
}Variable explorer search & filter all 8 variables
Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.
| Variable | Type | Distribution | Label | Definition | Units | In files | Source |
|---|---|---|---|---|---|---|---|
female_share# | continuous | Share of female students (covariate) | Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate. | share (0-1) | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
gpa# | continuous | Average GPA of low-income students (outcome) | The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file). | GPA points | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
id# | identifier | – | School identifier (1-35) | High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program. | id | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open |
post# | dummy | Post-adoption period (1 = yes) | 1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file. | 0/1 | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
time# | continuous | Period (2x2: 1-2; event study: 1-8) | Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5. | period | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
timeToTreat# | continuous | Periods relative to adoption (treated schools only) | Event time for treated schools: time - 5, from -4 to 3 (0 = first treated period, -1 = the reference period in the post). Missing for comparison schools, which are never treated. | periods | tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
treated# | dummy | Treatment group: school adopts tutoring (1 = yes) | 1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school. | 0/1 | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open | |
txp# | dummy | Treated x post: school is under the program (1 = yes) | The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate. | 0/1 | tutoring_did, tutoring_didevent | Corral and Yang (2024), via quarcs-lab/data-open |
Cross-file variable index
Which file each variable appears in (● = present).
| Variable | tutoring_did | tutoring_didevent |
|---|---|---|
female_share | ● | ● |
gpa | ● | ● |
id | ● | ● |
post | ● | ● |
time | ● | ● |
timeToTreat | ● | |
treated | ● | ● |
txp | ● | ● |
Construction & formulas
Design variables
treated= 1 for the 10 schools that adopt the program (id26–35), 0 for the 25 comparison schools. Fixed within school.post= 1 after adoption:time= 2 in the 2×2 file;time≥ 5 in the event-study file.txp=treated×post: the DiD treatment indicator (1 for a treated school in a treated period).timeToTreat=time− 5 for treated schools (−4, …, 3; 0 = first treated period); missing for comparison schools, which are never treated.
Estimates the post computes from these files
- 2×2 DiD: (mean treated post − mean treated pre) − (mean comparison post − mean comparison pre) = 36.20 − 10.886 = 25.315 (25.32 from means rounded to two decimals).
- Two-way fixed effects:
gpa ~ txp | id + time, standard errors clustered by school: 25.315 (SE 0.585). - Event study:
gpa ~ i(timeToTreat, ref = -1) | id + timeon the 280-row file, with comparison schools' missingtimeToTreatreplaced by a placeholder so they stay in the sample.
The datasets
Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.
expand to search (Ctrl/⌘+F) or print across all datasets
Variable dictionary
| Variable | Label | Definition | Construction | Units | Source | Coverage |
|---|---|---|---|---|---|---|
id identifier | School identifier (1-35) | High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program. | As provided in the original .dta files. | id | Corral and Yang (2024), via quarcs-lab/data-open | 35 schools |
time continuous | Period (2x2: 1-2; event study: 1-8) | Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5. | As provided in the original .dta files. | period | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
treated dummy | Treatment group: school adopts tutoring (1 = yes) | 1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school. | 1 if id is 26-35. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
post dummy | Post-adoption period (1 = yes) | 1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file. | post = (time == 2) in the 2x2 file; post = (time >= 5) in the event-study file. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
txp dummy | Treated x post: school is under the program (1 = yes) | The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate. | txp = treated x post. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
gpa continuous | Average GPA of low-income students (outcome) | The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file). | Simulated by Corral and Yang (2024); stored as single-precision float in the original .dta, written at full precision in the CSV. | GPA points | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
female_share continuous | Share of female students (covariate) | Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate. | Simulated by Corral and Yang (2024); single-precision float in the original .dta. | share (0-1) | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
Distribution & statistics (click a header to sort)
| Variable | Distribution | Coverage | N | Distinct | Min | Mean | Median | Max | SD |
|---|---|---|---|---|---|---|---|---|---|
id | – | 100% | 70 | 35 | — | — | — | — | — |
time | 100% | 70 | 2 | 1.00 | 1.50 | 1.50 | 2.00 | 0.504 | |
treated | 100% | 70 | 2 | 0 | 0.286 | 0 | 1.00 | 0.455 | |
post | 100% | 70 | 2 | 0 | 0.500 | 0.500 | 1.00 | 0.504 | |
txp | 100% | 70 | 2 | 0 | 0.143 | 0 | 1.00 | 0.352 | |
gpa | 100% | 70 | 70 | 59.39 | 77.12 | 76.27 | 99.15 | 10.88 | |
female_share | 100% | 70 | 70 | 0.471 | 0.528 | 0.527 | 0.570 | 0.027 |
Variable dictionary
| Variable | Label | Definition | Construction | Units | Source | Coverage |
|---|---|---|---|---|---|---|
id identifier | School identifier (1-35) | High-school identifier, 1 to 35. Schools 26-35 adopt the tutoring program. | As provided in the original .dta files. | id | Corral and Yang (2024), via quarcs-lab/data-open | 35 schools |
time continuous | Period (2x2: 1-2; event study: 1-8) | Time period. In the 2x2 file, 1 = before and 2 = after the program; in the event-study file, periods 1-8 with adoption at the start of period 5. | As provided in the original .dta files. | period | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
treated dummy | Treatment group: school adopts tutoring (1 = yes) | 1 for the 10 schools that adopt the after-school tutoring program, 0 for the 25 comparison schools. Constant within school. | 1 if id is 26-35. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
gpa continuous | Average GPA of low-income students (outcome) | The outcome: the school's average GPA of low-income students, described as a 0-100 score (simulated values reach 107.68 in the event-study file). | Simulated by Corral and Yang (2024); stored as single-precision float in the original .dta, written at full precision in the CSV. | GPA points | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
female_share continuous | Share of female students (covariate) | Share of female students in the school, between 0 and 1. Varies within schools over time; used as a robustness covariate. | Simulated by Corral and Yang (2024); single-precision float in the original .dta. | share (0-1) | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
post dummy | Post-adoption period (1 = yes) | 1 in periods after the program starts: time = 2 in the 2x2 file, time >= 5 in the event-study file. | post = (time == 2) in the 2x2 file; post = (time >= 5) in the event-study file. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
txp dummy | Treated x post: school is under the program (1 = yes) | The DiD treatment indicator: 1 for a treated school in a post-adoption period. Its coefficient in the regressions is the DiD estimate. | txp = treated x post. | 0/1 | Corral and Yang (2024), via quarcs-lab/data-open | all rows |
timeToTreat continuous | Periods relative to adoption (treated schools only) | Event time for treated schools: time - 5, from -4 to 3 (0 = first treated period, -1 = the reference period in the post). Missing for comparison schools, which are never treated. | timeToTreat = time - 5 if treated == 1; missing otherwise. | periods | Corral and Yang (2024), via quarcs-lab/data-open | 80 of 280 rows (treated schools) |
Distribution & statistics (click a header to sort)
| Variable | Distribution | Coverage | N | Distinct | Min | Mean | Median | Max | SD |
|---|---|---|---|---|---|---|---|---|---|
id | – | 100% | 280 | 35 | — | — | — | — | — |
time | 100% | 280 | 8 | 1.00 | 4.50 | 4.50 | 8.00 | 2.30 | |
treated | 100% | 280 | 2 | 0 | 0.286 | 0 | 1.00 | 0.453 | |
gpa | 100% | 280 | 280 | 60.08 | 80.14 | 78.53 | 107.7 | 12.20 | |
female_share | 100% | 280 | 280 | 0.470 | 0.521 | 0.524 | 0.570 | 0.028 | |
post | 100% | 280 | 2 | 0 | 0.500 | 0.500 | 1.00 | 0.501 | |
txp | 100% | 280 | 2 | 0 | 0.143 | 0 | 1.00 | 0.351 | |
timeToTreat | 29% | 80 | 8 | -4.00 | -0.500 | -0.500 | 3.00 | 2.31 |
Known limitations & caveats
- Simulated, not observed. The schools come from the teaching example of Corral and Yang (2024), built to illustrate DiD mechanics. The R² near 0.99 and a 25-point effect on GPA are far larger than real tutoring programs produce.
- GPA can exceed 100.
gpais described as a 0–100 score, but the simulated event-study values reach 107.68 (the 2×2 file stays below 100, maximum 99.15). Treat it as a continuous score, not a bounded grade. - timeToTreat is missing by design for the 25 comparison schools. Software that drops rows with missing values (pyfixest, fixest, Stata's
i.operator) silently removes every comparison school; fill it with a placeholder or build the event dummies by hand. - Two independent simulations. The 2×2 file is not periods 4 and 5 of the event-study file, so estimates from the two files differ (25.315 vs. a collapsed 24.897).
- female_share varies within schools over time. It is a time-varying covariate; the post uses it only as a robustness check (it moves the DiD estimate by 0.013).
- CSV values carry full float precision. The original .dta files store
gpaandfemale_shareas single-precision floats; the CSVs write those exact values as doubles (17 significant digits). In Stata, read the CSV withimport delimited …, asdouble.