Synthetic Control in Stata: Interactive Lab

A companion to The Synthetic Control Method in Stata: Did Proposition 99 Cut Smoking in California? ↗ Back to the post

Did Proposition 99 reduce cigarette sales in California?

Proposition 99 took effect in California in January 1989. It raised the state cigarette tax by 25 cents per pack and earmarked the new revenue for health and anti-smoking programs. The causal question asks what cigarette sales in California would have been without the policy. No data record that path, so we must estimate it. The synthetic control method builds the missing path as a weighted average of other states that matched California before 1989.

This app presents the main results of the Stata post in four tabs. The first tab defines the key terms, and the second shows the five donor states that build synthetic California. The third tab traces the gap between actual and synthetic sales, which reaches −26.37 packs per capita in 1999. The fourth tab ranks California against placebo runs for the other 38 states. All estimates come from the Stata log, except the pre-1989 synthetic path, which we recompute from the data with the same rounded weights.

The donor recipe in one picture

Synthetic California is built from only five of the 38 donor states. The other 33 states receive a reported weight of zero. Sparse weights like these are typical of synthetic control, because the weights must be nonnegative and sum to one.

The five weights are 33.4% for Utah, 23.5% for Nevada, 20.2% for Montana, 16.1% for Colorado, and 6.8% for Connecticut. The optimizer selects these states because their weighted average matches the predictors of California before 1989. No single donor needs to resemble California on its own.

Tab 2

Donor Recipe

The chart shows the weights of all 38 donor states. The table compares the predictors of California with those of its synthetic version. A final note explains why the V-weights are fragile.

Tab 3

The Gap

The chart follows actual and synthetic sales from 1970 to 2000. A switch isolates the gap that opens after 1989. Four readouts report the average effect, the peak gap, the gap in 2000, and the pre-1989 fit.

Tab 4

Placebo Ranking

The chart ranks California against 38 placebo runs. A filter drops poorly fitted placebos with the cut(2) rule. California ranks first in both views.

Glossary (open a card if a term is unfamiliar)

Synthetic control method (SCM)
The method builds a weighted average of untreated units that reproduces the pre-treatment path of the treated unit. After treatment, this synthetic unit stands in for the missing outcome without the policy. Abadie and Gardeazabal (2003) introduced the method, and Abadie, Diamond, and Hainmueller (2010) applied it to Proposition 99.
Donor pool
The donor pool holds the 38 states other than California. Abadie, Diamond, and Hainmueller (2010) built the data without states that had large tobacco control programs or raised cigarette taxes by 50 cents or more over 1989–2000. They also left out the District of Columbia, and the analysis code removes no further states.
Unit weights W
The unit weights are nonnegative and sum to one. The inner optimization chooses them to bring the predictors of synthetic California close to those of California. Utah receives 33.4%, Nevada 23.5%, Montana 20.2%, Colorado 16.1%, and Connecticut 6.8%, while the other 33 states receive zero.
Predictor weights V
The diagonal matrix V sets how much each predictor counts in the inner problem. The outer loop picks V to minimize the pre-treatment MSPE of the outcome. In the baseline run, age15to24 (0.546) and cigsale(1975) (0.422) dominate. These weights are poorly identified: the placebo refit of California gives 0.002 and 0.768 with almost the same donor weights.
Pre-treatment fit (RMSE, R²)
These statistics measure how closely the synthetic path follows California before 1989. The baseline fit has an RMSE of 1.756 packs, and synth2 reports an R² of 0.974. Its R² formula divides by the variation of the synthetic series, while the conventional R² with the same weights is 0.976. The fit is close but not exact, since the largest pre-treatment gap is 5.88 packs, in 1970. A good fit is necessary for a credible counterfactual, but it is not sufficient.
ATT (average treatment effect on the treated)
Each yearly gap between actual and synthetic sales is an effect estimate for that year. The ATT is the average of these gaps over 1989–2000, which equals −19.00 packs per capita. The yearly gaps range from −7.59 packs in 1989 to −26.37 packs in 1999.
In-space placebo
The test reruns the method with each donor state as if it had been treated. California ranks first among 39 units by the ratio of post-treatment to pre-treatment MSPE. It also ranks first among the 20 units that pass the cut(2) fit filter.
MSPE ratio (post/pre)
The ratio divides the post-treatment MSPE of a unit by its pre-treatment MSPE. A large ratio means a large gap after 1989 relative to the fit before 1989. California has a ratio of 123.5, followed by Georgia (80.0), Virginia (79.0), and Missouri (70.9). The ratio of California uses the pre-MSPE of the placebo refit, 3.17, rather than that of the baseline fit.

The donor recipe: which states build synthetic California?

The optimization searches the 38-state donor pool for the convex combination that best matches the predictors of California. The solution is sparse: only five states receive positive weight. The chart below sorts all 38 donor weights, and the table compares the seven predictors.

Optimal donor weights

Steel blue marks the largest donor, Utah. Orange marks the other four donors with positive weight: Nevada, Montana, Colorado, and Connecticut. The remaining 33 states appear dimmed because their reported weights are zero.

Predictor balance: why these five states?

For each of the seven predictors, the table compares California, synthetic California, and the simple average of the 38 donor states. The percentages in parentheses are the differences from California that synth2 reports in its Bias columns. Orange marks the donor average in the one row where it misses California by more than 20%.

Consider cigsale(1988), where California sold 90.1 packs per capita and synthetic California sold 91.7. The simple donor average of 113.8 misses by 26.3%, while the synthetic value misses by only 1.7%. In 1975 the synthetic value of 127.1 sits close to California, which is expected because cigsale(1975) carries a large V-weight. A close match on the predictors is necessary for a credible counterfactual, but it does not prove that the counterfactual is right.

What to look for

  • Sparse weights are a feature of the method. The weights must be nonnegative and sum to one, so the solution usually rests on a few donors. Here, 33 of the 38 states receive zero weight. A difference-in-differences design, by contrast, gives every control state the same weight.
  • The predictor match is close, except for income. Six of the seven differences are at most 1.74% in absolute value, and five are at most 0.3%. Log GDP per capita is the exception. Its −2.16% difference compares logarithms, so synthetic California falls short by 0.22 log points, a GDP per capita about 20% lower. The simple donor average misses cigsale(1988) by 26.3% and cigsale(1980) by 14.9%.
  • The V-weights are fragile. In the baseline run, age15to24 (0.546) and cigsale(1975) (0.422) carry 97% of the V mass. The placebo refit of California gives them 0.002 and 0.768, while the donor weights barely move. We therefore should not read V as a ranking of which predictors matter.

Actual and synthetic California, 1970–2000

The chart plots cigarette sales in California and in its synthetic version. The shaded area marks the pre-treatment years 1970–1988, when the two paths should stay close. After 1989, the gap between them estimates the effect of Proposition 99 in packs per capita. The log prints synthetic values for 1989–2000 only, so we recompute the 1970–1988 values from the data with the rounded Stata weights.

View

Average ATT (1989–2000)
−19.00
packs per capita per year
Peak gap (1999)
−26.37
largest single-year gap
Gap in 2000
−25.76
38.2% below the synthetic path
Pre-1989 RMSE
1.756
R² 0.974 (synth2 definition)

What to look for

  • Before 1989, the paths are close but not identical. Actual sales rise from 123.0 packs in 1970 to 128.0 in 1976 and then fall to 90.1 packs in 1988. The synthetic path stays close in every year except 1970, when the gap is 5.88 packs. From 1971 onward, the largest gap in absolute value is 2.24 packs, in 1987. The RMSE of 1.756 packs summarizes this fit.
  • After 1989, the paths diverge. The gap opens at −7.59 packs in 1989 and widens through the 1990s, with brief narrowings in 1995, 1998, and 2000. It reaches its largest value, −26.37 packs, in 1999. The gap of −25.76 packs in 2000 equals 38.2% of the synthetic value, and the share in 1999 is 35.8%.
  • The effect grows rather than fades. The gap widens over most of the post-treatment period. Because Proposition 99 combined a tax increase with anti-smoking spending, the design cannot identify which component drives this growth. The gaps of 1999 and 2000 also coincide with Proposition 10, which raised the state tax by a further 50 cents per pack in January 1999.

In-space placebo ranking: how unusual is the California gap?

The in-space placebo test reruns the method with each of the other 38 states as the treated unit. If Proposition 99 had no effect, the MSPE ratio of California should look ordinary in this distribution. The chart below shows where it actually falls.

Filter

In the full view, dimmed bars mark the 19 states that the cut(2) rule removes.
California ratio (post/pre MSPE)
123.5
pre 3.17, post 391.25 (placebo refit)
Rank (trimmed)
1 of 20
p = 0.050
Rank (full)
1 of 39
p = 0.026
Runner-up ratio (Georgia)
80.0
California exceeds it by 43.5

How to read this

  • Switch to the full view to see all 39 units. California has the highest MSPE ratio, 123.5, which is 54% larger than the ratio of Georgia (80.0). If the policy had been assigned to one of the 39 states at random, the chance of drawing a ratio this large would be 1/39 ≈ 0.026.
  • Switch back to the trimmed view. The cut(2) rule drops the 19 states whose pre-treatment MSPE exceeds twice that of California. A high pre-MSPE sits in the denominator, so it makes the ratio small, as for New Hampshire, whose ratio is below 0.1. All 19 dropped states have ratios below that of California, and the highest, for Indiana, is 32.6. Dropping them therefore cannot change the rank of California. The cut only shrinks the reference set from 39 to 20 units, which raises the smallest attainable p-value from 0.026 to 0.050. It matters more for the yearly gap comparisons in the post, where poorly fitted placebos would add noise. Among the 20 remaining units, California again ranks first, so p = 1/20 = 0.050.
  • Read the two p-values together. The full ranking gives p = 0.026, and the trimmed ranking gives p = 0.050. With 20 units, 0.050 is the smallest p-value that the test can produce. Both rankings place California first, which supports a real effect of Proposition 99 on cigarette sales.

Connecting back to the previous tabs

The placebo ranking and the donor weights should be read together. The five donors of synthetic California do not rank near the top of the placebo distribution. Georgia, Virginia, Missouri, Texas, and Oklahoma do, because their post-treatment gaps are large relative to their pre-treatment fit. Four of the five donors, namely Utah, Nevada, Colorado, and Connecticut, are themselves removed by cut(2) because their own fits are poor. Good donors can therefore be poor placebos. Utah and Nevada lie near the extremes of the pool, so blends of other states fit them poorly, yet together they bracket California.

The post also runs an in-time placebo with the fake treatment year 1985. That run yields gaps of −3.33 to −8.65 packs over 1985–1988, much smaller than the gaps of −13.97 to −25.59 packs after 1989. The leave-one-out test drops each of the five donors with positive weight in turn. Across those five refits, the gap in 2000 ranges from −28.35 to −23.49 packs and never approaches zero. Both checks agree with the baseline estimate, although neither can rule out every alternative explanation.