Downloads
Each dataset is available as a labeled Stata .dta and its source file.
⇩ Download all data (ZIP)stata_codebook.do
| Dataset | Grain | Rows | Stata | Source |
|---|---|---|---|---|
smoking_sc | state × year (balanced panel) | 1,209 × 7 | smoking_sc.dta | smoking_sc.csv |
Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.
Load directly in code
Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.
Stata
* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
use "${BASE}smoking_sc.dta", clear
describe
notesPython
!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
df = pd.read_stata(BASE + "smoking_sc.dta")
# load every dataset at once
files = ["smoking_sc"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}
# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "smoking_sc.dta", "smoking_sc.dta")
df, meta = pyreadstat.read_dta("smoking_sc.dta")Copy and paste this snippet into an empty Google Colab notebook: https://colab.research.google.com/notebooks/empty.ipynb
R
# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
df <- read_dta(paste0(BASE, "smoking_sc.dta"))Overview & sources
Proposition 99 raised the cigarette tax in California by 25 cents per pack from January 1989 and funded anti-smoking education. Abadie, Diamond, and Hainmueller (2010), hereafter ADH (2010), evaluate it with a panel of 39 US states from 1970 to 2000. Each row is one state in one year, so the panel has 1,209 rows. California is the treated state, and the other 38 states form the donor pool. The outcome cigsale measures cigarette sales per capita, in packs. Four covariates describe each state: log GDP per capita, the share of the population aged 15–24, the retail price of cigarettes, and beer consumption per capita. The Python tutorial builds a synthetic California from these variables. It estimates that the program reduced sales by 18.98 packs per capita per year over 1989–2000.
smoking_sc has 39 states × 31 years = 1,209 rows, keyed by state × year and sorted by state and year. The sample omits four states that began large tobacco control programs in 1989–2000 (Arizona, Florida, Massachusetts, and Oregon). It also omits seven states that raised cigarette taxes by 50 cents or more over the same period, as well as the District of Columbia (Section 3.2 of ADH 2010). The original Stata file stores state as a numeric code with value labels in alphabetical order, so California has code 3. The CSV copy stores the state name as text instead, and so does the .dta file generated from it in this folder.
Data sources
| Source | Provides | Reference / URL |
|---|---|---|
| Abadie, Diamond, and Hainmueller (2010) | The case study and the state panel. Appendix A names the original sources: Orzechowski and Walker (2005) for sales and prices, the Bureau of the Census for income and the age share, and the Beer Institute for beer consumption. | Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746 |
| QuaRCS-lab data-open | The original Stata file smoking_sc.dta, labeled Tobacco Sales in 39 US States, which the post and its Stata edition load | https://github.com/quarcs-lab/data-open/raw/master/isds/smoking_sc.dta |
| This post (CSV copy) | smoking_sc.csv, with the state name as text and the values of the original .dta file at full double precision (89.8 appears as 89.80000305175781). The post loads this copy first and falls back to the original file online. | Mendez, C. (2026). https://carlos-mendez.org/post/python_sc101/ |
Cite this data
Please cite this dataset as follows.
APA
Mendez, C. (2026). Proposition 99 and cigarette sales in 39 US states (synthetic control tutorial) [Data set]. https://carlos-mendez.org/post/python_sc101/
Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of California's tobacco control program. Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746BibTeX
@misc{mendez2026pythonsc101,
author = {Mendez, Carlos},
title = {Proposition 99 and cigarette sales in 39 US states (synthetic control tutorial)},
year = {2026},
howpublished = {\url{https://carlos-mendez.org/post/python_sc101/}},
note = {Data set}
}
@article{abadie2010synthetic,
author = {Abadie, Alberto and Diamond, Alexis and Hainmueller, Jens},
title = {Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of {California's} Tobacco Control Program},
journal = {Journal of the American Statistical Association},
volume = {105}, number = {490}, pages = {493--505}, year = {2010},
doi = {10.1198/jasa.2009.ap08746}
}Variable explorer search & filter all 7 variables
Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.
| Variable | Type | Distribution | Label | Definition | Units | In files | Source |
|---|---|---|---|---|---|---|---|
age15to24# | continuous | Share of the population aged 15–24, a fraction (covariate) | Share of the state population aged 15 to 24, stored as a fraction between 0.129 and 0.204 (mean 0.175). The original label calls it a percent, and Table 1 of ADH (2010) reports it in percent, so multiply it by 100 to compare. | fraction (0–1) | smoking_sc | US Census Bureau, via ADH (2010) and quarcs-lab/data-open | |
beer# | continuous | Beer consumption per capita, in gallons (covariate) | Per capita consumption of malt beverages, in gallons. The data start in 1984, so the 1980–1988 average in the post rests on 1984–1988 alone. | gallons per capita | smoking_sc | Beer Institute, via ADH (2010) and quarcs-lab/data-open | |
cigsale# | continuous | Cigarette sales per capita, in packs (outcome) | Annual cigarette sales per capita, in packs, and the outcome of the analysis. The series divides the tax-paid sales of cigarette packs in a state by its population (Appendix A of ADH 2010). | packs per capita (annual) | smoking_sc | Orzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open | |
lnincome# | continuous | Log of state GDP per capita (covariate) | Natural logarithm of GDP per capita in the state, as the original .dta label and Table 1 of ADH (2010) describe it. The text and Appendix A of ADH (2010) call it per capita state personal income (logged), converted to 1997 dollars with the Consumer Price Index. | log of 1997 dollars per person | smoking_sc | Bureau of the Census, United States Statistical Abstract, via ADH (2010) and quarcs-lab/data-open | |
retprice# | continuous | Retail price of cigarettes, in cents per pack (covariate) | Average retail price of a pack of cigarettes, in cents, including state sales taxes where applicable. Appendix A of ADH (2010) converts income to 1997 dollars but mentions no such conversion for this price. | cents per pack | smoking_sc | Orzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open | |
state# | identifier | n/a | State name (39 US states) | Name of the US state. California is the treated state, and the other 38 states form the donor pool. | smoking_sc | ADH (2010), via quarcs-lab/data-open | |
year# | year | n/a | Year (1970–2000) | Calendar year of the observation. Proposition 99 takes effect in 1989, so 1970–1988 is the pre-treatment period and 1989–2000 the post-treatment period. | calendar year | smoking_sc | ADH (2010), via quarcs-lab/data-open |
Cross-file variable index
Which file each variable appears in (● = present).
Construction & formulas
How the post builds the predictors
- Covariate averages. For each state, the post averages
lnincome,age15to24,retprice, andbeerover 1980–1988. Missing years drop out of each average, so thebeeraverage rests on 1984–1988 alone. - Lagged sales. The columns
cigsale_1988,cigsale_1980, andcigsale_1975repeat the sales of each state in that year in every row of the state. Each one enters the fit as a predictor with its own one-year window. - Treatment indicator. The column
treatedequals 1 for California in 1989–2000, which gives 12 rows, and 0 otherwise.
What the post estimates from this file
- Synthetic California. It is a weighted average of donor sales, with nonnegative weights that sum to one. Five states receive positive weight: Utah (0.335), Nevada (0.236), Montana (0.202), Colorado (0.160), and Connecticut (0.068).
- Gap and ATT. The gap in a year is actual sales in California minus synthetic sales. The average treatment effect on the treated (ATT) is the mean gap over 1989–2000, −18.98 packs per capita per year.
- Pre-treatment fit. The root mean squared error of the gaps over 1970–1988 is 1.754 packs per capita.
The datasets
Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.
expand to search (Ctrl/⌘+F) or print across all datasets
Variable dictionary
| Variable | Label | Definition | Construction | Units | Source | Coverage |
|---|---|---|---|---|---|---|
state identifier | State name (39 US states) | Name of the US state. California is the treated state, and the other 38 states form the donor pool. | The original .dta file stores a numeric code with value labels in alphabetical order (1 = Alabama, 3 = California, 39 = Wyoming). The CSV keeps the label text. | ADH (2010), via quarcs-lab/data-open | 39 states in every year | |
year year | Year (1970–2000) | Calendar year of the observation. Proposition 99 takes effect in 1989, so 1970–1988 is the pre-treatment period and 1989–2000 the post-treatment period. | Stored as a float in the original .dta file and as an integer in the CSV. | calendar year | ADH (2010), via quarcs-lab/data-open | 31 years for every state |
cigsale continuous | Cigarette sales per capita, in packs (outcome) | Annual cigarette sales per capita, in packs, and the outcome of the analysis. The series divides the tax-paid sales of cigarette packs in a state by its population (Appendix A of ADH 2010). | Original .dta label: cigarette sale per capita (in packs). The post compares California with its synthetic control on this variable in every year. | packs per capita (annual) | Orzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open | 1970–2000, all 1,209 rows |
lnincome continuous | Log of state GDP per capita (covariate) | Natural logarithm of GDP per capita in the state, as the original .dta label and Table 1 of ADH (2010) describe it. The text and Appendix A of ADH (2010) call it per capita state personal income (logged), converted to 1997 dollars with the Consumer Price Index. | Original .dta label: log state per capita gdp. The post averages it over 1980–1988 as a predictor. | log of 1997 dollars per person | Bureau of the Census, United States Statistical Abstract, via ADH (2010) and quarcs-lab/data-open | 1972–1997 (1,014 of 1,209 rows) |
beer continuous | Beer consumption per capita, in gallons (covariate) | Per capita consumption of malt beverages, in gallons. The data start in 1984, so the 1980–1988 average in the post rests on 1984–1988 alone. | Original .dta label: beer consumption per capita. ADH (2010) also average it over 1984–1988 (notes to Table 1). | gallons per capita | Beer Institute, via ADH (2010) and quarcs-lab/data-open | 1984–1997 (546 of 1,209 rows) |
age15to24 continuous | Share of the population aged 15–24, a fraction (covariate) | Share of the state population aged 15 to 24, stored as a fraction between 0.129 and 0.204 (mean 0.175). The original label calls it a percent, and Table 1 of ADH (2010) reports it in percent, so multiply it by 100 to compare. | Original .dta label: percent of state population aged 15-24 years. The post averages it over 1980–1988 as a predictor. | fraction (0–1) | US Census Bureau, via ADH (2010) and quarcs-lab/data-open | 1970–1990 (819 of 1,209 rows) |
retprice continuous | Retail price of cigarettes, in cents per pack (covariate) | Average retail price of a pack of cigarettes, in cents, including state sales taxes where applicable. Appendix A of ADH (2010) converts income to 1997 dollars but mentions no such conversion for this price. | Original .dta label: retail price of cigarettes. The post averages it over 1980–1988 as a predictor. | cents per pack | Orzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open | 1970–2000, all 1,209 rows |
Distribution & statistics (click a header to sort)
| Variable | Distribution | Coverage | N | Distinct | Min | Mean | Median | Max | SD |
|---|---|---|---|---|---|---|---|---|---|
state | n/a | 100% | 1,209 | 39 | n/a | n/a | n/a | n/a | n/a |
year | n/a | 100% | 1,209 | 31 | 1970 | 1985.0 | 1985 | 2000 | 8.95 |
cigsale | 100% | 1,209 | 703 | 40.70 | 118.9 | 116.3 | 296.2 | 32.77 | |
lnincome | 84% | 1,014 | 1,014 | 9.40 | 9.86 | 9.86 | 10.49 | 0.171 | |
beer | 45% | 546 | 145 | 2.50 | 23.43 | 23.30 | 40.40 | 4.22 | |
age15to24 | 68% | 819 | 819 | 0.129 | 0.175 | 0.178 | 0.204 | 0.015 | |
retprice | 100% | 1,209 | 849 | 27.30 | 108.3 | 95.50 | 351.2 | 64.38 |
Known limitations & caveats
- The variable
lnincomemeasures GDP, not income. Despite its name,lnincomeis the log of GDP per capita, according to the original label and Table 1 of ADH (2010). The text and Appendix A of the same paper call it per capita state personal income (logged), so the source itself uses both descriptions. - The variable
age15to24is a share, not a percent. Its values lie between 0.129 and 0.204, although the original label reads percent. Multiply it by 100 to compare it with the percentages in Table 1 of ADH (2010). - Covariate coverage is incomplete. The variable
lnincomecovers 1972–1997,beercovers 1984–1997, andage15to24covers 1970–1990. Thebeeraverage over 1980–1988 therefore rests on 1984–1988 alone. A fake start in 1984 leaves no beer data in its window of 1980–1983, so that in-time placebo fails in the post. - State coding differs across files. The original .dta file stores
stateas a numeric code with value labels, and the CSV and the.dtafile generated here store the state name as text. In Stata,encode state, generate(state_id)restores the original codes, becauseencodenumbers the names in alphabetical order. California then has code 3, as in the original file. - Read the CSV at full precision. The original file stores every number in single precision, and the CSV writes these exact values as doubles. Some parsers miss the last digits. The default parser of pandas 3.0.1 changes 1,422 of the 4,797 numeric values, and
readr::read_csv2.1.6 changes 771, each by less than 1e-13. The pandas optionfloat_precision="round_trip"recovers the exact values, andload_data()in the post uses it. The functionsread.csvin base R andimport delimitedin Stata 19 also read them exactly. - The generated .dta file stores doubles. The renderer writes every number through pyreadstat as a double, including
year, whereas the original file stores floats. The values themselves are identical to those of the original file. In Stata,recast float cigsale lnincome beer age15to24 retpricerestores the original storage type without changing any value. - Sales come from tax records. Cigarette sales per capita rest on state tax revenues rather than on surveys of smoking. ADH (2010) note that smuggling across tax jurisdictions affects such data (Section 3.2).
- A second tax increase in 1999. Proposition 10 raised the cigarette tax in California by a further 50 cents per pack in January 1999. Sales in California in 1999 and 2000 therefore reflect both measures.