The Synthetic Control Method in Stata

Did Proposition 99 cut smoking in California?

−19.00ATT · packs per capita / yr
0.974pre-treatment R² (synth2)
0.026in-space placebo p-value

Carlos Mendez

Nagoya University (GSID)

October 5, 2026

The Tension

Act I

Smoking was falling everywhere, so did the tax change anything?

In 1988, California voters passed Proposition 99, effective January 1989. It raised the cigarette tax by 25 cents per pack. It also funded anti-smoking education.

Yet cigarette sales were already falling nationwide. In California, the decline began in the late 1970s. What would California have done without the law?

A raw comparison hints at an effect, but the comparator is crude

Cigarette sales per capita, 1970–2000: California (solid blue) versus the unweighted average of 38 control states (dashed gray). The orange line marks Proposition 99 (1989).

Where we are going

  • The lab: a 39-state, 31-year panel and the ATT estimand
  • Build “synthetic California” with synth2, a weighted blend of donor states
  • Three inference tools: in-space placebo, in-time placebo, and leave-one-out
  • The lesson: prefer a transparent counterfactual to an untestable parallel-trends assumption

The Investigation

Act II

The lab: 39 states × 31 years, one treated unit, the ATT

  • Outcome: cigsale, cigarette sales per capita (packs)
  • Treatment: Proposition 99, California only, from 1989
  • Predictors: log GDP per capita, share aged 15–24, retail price, beer, and cigsale in 1975, 1980, and 1988

Strongly balanced panel: 1,209 observations, 19 pre-treatment years, and 12 post-treatment years. The estimand is the ATT for California, the one treated unit, not the ATE.

SCM builds a counterfactual by matching pre-treatment predictors

\[\min_{W} \sum_{m=1}^{M} v_m \left( X_{1m} - \sum_{j=2}^{J+1} w_j X_{jm} \right)^2\]

We pick donor weights \(w_j \ge 0\) that sum to one and match the predictors \(X_{1m}\) of California. The weights \(v_m\) set how much each predictor matters. The solution is the closest feasible blend of donors.

The synthetic control is a convex combination of real states, so it never extrapolates beyond the donor pool.

The treatment effect is the gap between actual and synthetic California

\[\hat{\tau}_t = Y_{1t} - \sum_{j=2}^{J+1} w_j^* Y_{jt}\]

The effect in year \(t\) is actual minus synthetic sales. A negative \(\hat{\tau}_t\) means that the law lowered sales. Each year from 1989 to 2000 yields one such gap.

Identification rests on good pre-treatment fit, no interference, no anticipation, and no donor contamination; only the fit is directly visible in the data.

One synth2 call fits the baseline synthetic control

synth2 cigsale lnincome age15to24 retprice beer ///
    cigsale(1988) cigsale(1980) cigsale(1975), ///
    trunit(3) trperiod(1989) xperiod(1980(1)1988) nested allopt

trunit(3) = California · trperiod(1989) = treatment year · nested = outer V / inner W · allopt = multiple starts to avoid local optima.

Synthetic California tracks the pre-1989 path closely, but not exactly

Actual versus synthetic California, 1970–2000: close before 1989 apart from 1970, then a widening gap.

Age 15–24 and 1975 sales dominate the V-weights, but V is fragile

Predictor (V-matrix) weights: how much each predictor counts in the matching distance.

Synthetic California is a blend of five states, led by Utah at one-third

Donor weights: the five states that compose synthetic California (33 others get a reported weight of zero).

By 2000, real California sold 38% fewer packs than its synthetic twin

−19.00

average ATT, packs per capita / yr (1989–2000); −7.59 in 1989 deepening to −26.37 by 1999

The gap widens through the 1990s, though not in every year

Treatment effect (actual minus synthetic California) over time; the negative gap widens after 1989.

The Resolution

Act III

California has the most extreme signal-to-noise ratio of all 39 states

State Pre MSPE Post/Pre MSPE
California 3.17 123.5
Georgia 1.46 80.0
Virginia 2.78 79.0
Missouri 1.20 70.9

The post/pre MSPE ratio asks how much worse the fit becomes after 1989 than before. California has the highest ratio, well above the runner-up.

Run the placebo on every state: California is the clear outlier

Gaps for California (purple) and the 19 placebo states retained by cut(2): the lines overlap near zero before 1989, and California falls below almost every placebo afterward.

Across all 39 states, a ratio this large occurs 2.6% of the time

0.026

in-space placebo p-value, all controls (1/39); p = 0.050 after the cut(2) fit filter (1/20)

Significance holds in most post-treatment years

Left-sided permutation p-values over time (left-sided, because the effect is negative): p = 0.050 in most years.

A fake 1985 treatment produces much smaller gaps

In-time placebo: actual versus synthetic California with a fake treatment in 1985. A modest gap opens during 1985–1988, and the lines split much further after the real 1989 policy.

The gap stays modest at the fake date and grows after the real one

In-time placebo effect: smaller gaps during the fake 1985–1988 window, larger gaps after the real 1989 treatment.

No single donor drives the sign of the effect: leave-one-out gaps stay negative

Leave-one-out output in seven panels. In the last two, each gray line drops one weighted donor, and the paths stay close to the full-pool fit.

Does choosing controls by fit make this causal? Not by itself

Objection. A data-driven counterfactual is still a model. It cannot manufacture identification. Pre-1989 fit does not guarantee post-1989 validity.

Response. Correct. The ATT is credible only under no interference, no anticipation, no donor contamination, and good pre-treatment fit. SCM makes the last condition visible and disciplines the comparator. It cannot rule out cross-border shopping or an idiosyncratic donor such as Utah. It also gives no standard errors, so inference rests on the placebos.