Canonical regional panel data-generating process¶
The canonical simulation represents a set of regions observed over repeated time periods before and after a pricing or promotion policy is introduced.
Its purpose is not to make synthetic data look realistic at any cost. The purpose is to create controlled causal experiments in which the true counterfactual and treatment effect are known.
Structural design¶
For region (i) and period (t), the untreated outcome is generated from
[ Y_{it}(0) = mu + alpha_i + gamma_t + s_t + eta_X X_i + eta_U U_i + arepsilon_{it}. ]
Here:
- (alpha_i) is persistent regional heterogeneity;
- (gamma_t) contains the common trend and common time shocks;
- (s_t) is deterministic seasonality;
- (X_i) is the observed regional confounder exposed as
market_pressure; - (U_i) is an unobserved demand factor;
- (arepsilon_{it}) is idiosyncratic outcome noise.
Policy assignment is intentionally non-random. Regions with larger values of
[ S_i = lambda_X X_i + lambda_U U_i + eta_i ]
are more likely to receive the policy. The simulator assigns the policy to the highest assignment scores. Because both (X_i) and (U_i) also affect the untreated outcome, naive treated-versus-control comparisons are confounded by construction.
Treatment-effect heterogeneity¶
Each region has a structural one-period treatment effect
[ au_i = au_0 + h(0.6 X_i - 0.4 U_i), ]
where (h) controls the strength of heterogeneity. The true population ATE and the ATT among policy-assigned regions are therefore known exactly.
Lag and carry-over¶
The observed policy path is converted into an effective exposure process. With lag (L) and carry-over parameter ( ho),
[ E_{it} = (1- ho)D_{i,t-L} +
ho E_{i,t-1}, qquad 0 leq ho < 1. ]
When (
ho=0), exposure is simply the lagged treatment indicator. Positive
values produce gradual adjustment and post-treatment persistence. A finite
treatment_duration can be used to create a policy pulse whose effect
continues after treatment has ended.
The realised causal effect is
[ Delta_{it} = au_i E_{it}, ]
and the potential outcome under the generated policy path is
[ Y_{it}( ext{policy}) = Y_{it}(0) + Delta_{it}. ]
Causal graph¶
graph LR
X[Observed regional factor X] --> A[Policy assignment]
U[Latent demand factor U] --> A
X --> Y0[Untreated outcome]
U --> Y0
R[Regional heterogeneity] --> Y0
T[Trend / seasonality / common shocks] --> Y0
A --> D[Observed treatment path]
D --> E[Effective exposure]
X --> H[Regional treatment effect]
U --> H
H --> Y[Observed outcome]
E --> Y
Y0 --> Y
This graph deliberately contains the back-door paths (A leftarrow X ightarrow Y(0)) and (A leftarrow U ightarrow Y(0)). Only the first is observable to an analyst.
Separation between observed data and causal truth¶
RegionalPanelData.observed contains only analyst-facing quantities such as
outcomes, treatment status, adoption timing, and the observed confounder.
RegionalPanelData.truth is kept separate and contains quantities that would
not be available in a real observational study:
- the latent confounder;
- assignment score;
- regional structural treatment effect;
- effective exposure;
- untreated potential outcome;
- policy-path potential outcome;
- exact realised causal effect.
That separation lets later estimators be evaluated against truth without accidentally giving the estimator information that would not exist in a real analysis.
Exact estimands¶
The simulator returns three exact reference quantities:
- ATE: mean structural one-period treatment effect across all regions;
- ATT: mean structural one-period treatment effect among policy-assigned regions;
- realized policy effect: mean realised causal effect over region-periods with positive effective exposure.
These quantities are population values for the generated synthetic sample, not estimates.