Difference-in-Differences and event studies¶
Difference-in-Differences (DiD) identifies treatment effects from differences in outcome changes rather than differences in outcome levels.
That distinction is useful when treated and comparison units have persistent level differences but would have evolved similarly in the absence of treatment.
Canonical 2x2 DiD¶
With treated group (G_i) and post-treatment indicator (P_t), the canonical regression is
[ Y_{it} = alpha + gamma G_i + lambda P_t + au(G_i P_t) + arepsilon_{it}. ]
The interaction coefficient ( au) equals
[ (ar Y_{T,post}-ar Y_{T,pre}) - (ar Y_{C,post}-ar Y_{C,pre}). ]
Under parallel trends, consistency, no anticipation, and the relevant no-interference condition, this identifies an ATT.
The repository exposes this design through estimate_did_2x2.
Two-way fixed-effects DiD¶
For multiple regions and periods,
[ Y_{it} = alpha_i + gamma_t + au D_{it} + arepsilon_{it}. ]
Entity effects absorb persistent regional level differences. Time effects absorb shocks common to all regions.
fit_twfe_did implements this baseline with entity-clustered uncertainty by
default.
For common treatment timing and homogeneous effects, it is a useful benchmark. It should not be interpreted as universally valid under staggered adoption.
Event-study specification¶
An event study replaces one treatment indicator with indicators for time relative to adoption:
[ Y_{it} = alpha_i + gamma_t + sum_{k eq -1} eta_k 1{t-G_i=k} + arepsilon_{it}. ]
The omitted event time is the reference period. By default it is (k=-1).
The implementation bins the minimum and maximum requested event times into tails. Never-treated regions have all event-time indicators equal to zero.
Post-treatment coefficients describe dynamic effects relative to the reference period under the identifying assumptions.
Pre-treatment coefficients¶
Lead coefficients are useful diagnostics. Under no anticipation and a stable parallel-trends relationship, they should not show systematic pre-treatment movement.
test_parallel_trends performs a joint Wald test that all estimated
pre-treatment event coefficients are zero.
A large p-value is not proof of parallel trends. Pretrend tests can have low power, and the identifying restriction concerns the unobserved post-treatment counterfactual.
Placebo intervention dates¶
estimate_placebo_did places a false intervention date entirely inside the
true pre-treatment sample.
A material placebo effect is evidence that the treatment and comparison groups were already evolving differently before treatment.
Placebos are falsification tools, not additional treatment-effect estimates.
Anticipation¶
No anticipation requires treatment not to affect outcomes before formal adoption.
If firms, customers, or regions react to an announced policy in advance, lead coefficients can move before event time zero.
The test suite includes a controlled violation in which treated regions receive an artificial pre-treatment outcome shift. The event study surfaces that shift in the lead coefficients.
Staggered adoption and treatment-effect heterogeneity¶
With staggered adoption, already-treated units can become implicit controls for later-treated units in a conventional TWFE regression.
When treatment effects differ across cohorts or over event time, the resulting TWFE coefficient can combine comparisons with non-obvious, and sometimes non-convex, weights.
The repository therefore includes estimate_group_time_att.
For cohort (g) and event time (k), it compares the cohort's change from (g-1) to (g+k) with the same change among regions that are:
- never treated; or
- not yet treated by (g+k).
Conceptually,
[ ATT(g,k) = E[Y_{g+k}-Y_{g-1}mid G=g] - E[Y_{g+k}-Y_{g-1}mid G>g+k ext{or never treated}]. ]
The implementation reports each cohort/event-time cell separately and a cohort-size-weighted average over the requested event times.
This design keeps heterogeneous cohort effects visible instead of forcing them into one coefficient.
Controlled heterogeneity failure case¶
The tests include two treated cohorts:
- an early cohort with effect 20;
- a later cohort with effect 0;
- never-treated controls.
The true average effect among treated regions is 10.
The group-time estimator recovers 10 exactly in the deterministic design. The single TWFE treatment coefficient differs because it mixes cohort comparisons.
That example is deliberately simple: its purpose is to make the weighting problem visible rather than to reproduce every modern staggered-DiD estimator.
What the diagnostics establish¶
The module deliberately separates estimates from diagnostics.
An event-study lead, a placebo estimate, or a pretrend test can reveal evidence against an identifying assumption.
They cannot establish that the identifying assumption is true.
The causal argument still depends on the design, the institutional setting, the comparison group, treatment timing, and the plausibility of the untreated counterfactual.