Case study: incremental effect of a regional pricing policy¶
Business question¶
What is the incremental effect of a regional pricing policy on weekly customer orders?
The case study uses the deterministic retail pricing benchmark introduced by the production workflow.
There are 48 regions observed for 84 weeks. Half of the regions adopt the pricing policy in week 52.
The analyst-facing dataset contains outcomes, policy timing, group membership, and an observed market-pressure covariate. Potential outcomes and the true treatment effect remain in a separate validation table.
Because this is a controlled benchmark, the true effect is known to be 6 incremental orders per treated region-week. That value is used only to validate the analysis after estimation.
Estimand¶
The primary estimand is the average incremental number of weekly orders per treated region during the post-policy period:
[ ATT = E[ Y_{it}(1)-Y_{it}(0) \mid G_i=1, t\ge 52 ]. ]
The business decision concerns what happened to policy regions because of the pricing intervention, not whether their observed order level is higher than comparison regions.
Why the raw group difference is not causal¶
Policy assignment is deliberately selective.
Policy regions differ systematically in persistent demand characteristics.
The raw post-policy contrast
[ E[Y\mid G=1,post] - E[Y\mid G=0,post] ]
therefore mixes:
- baseline regional differences;
- common time effects;
- the pricing-policy effect.
The generated case-study table reports this naive association beside the causal estimators so the distinction is visible.
Prediction is a different task¶
The case study also fits a factual-outcome prediction baseline.
It uses:
- time trend;
- sine/cosine seasonal terms;
- market pressure;
- region indicators;
- policy-group membership;
- observed policy status.
The model is evaluated on the final 16 weeks using a temporal holdout.
RMSE, MAE, and R-squared answer:
How accurately can we predict the observed weekly order volume?
They do not answer:
What would orders have been in policy regions if the pricing policy had not been introduced?
A high predictive R-squared therefore does not validate the causal estimate.
Identification strategy¶
The primary observational design is common-adoption Difference-in-Differences.
The two-way fixed-effects model is
[ Y_{it} = \alpha_i + \gamma_t + \tau D_{it} + \varepsilon_{it}. ]
Entity effects absorb persistent regional level differences.
Time effects absorb shocks common to all regions.
The causal interpretation requires:
- consistency;
- no anticipation;
- an appropriate no-interference condition;
- parallel untreated trends between policy and comparison regions.
The case study reports both the canonical 2x2 DiD and the panel TWFE estimate.
Diagnostics and falsification¶
The workflow includes:
- an event study around week 52;
- a joint pre-treatment coefficient test;
- placebo intervention dates at weeks 24, 32, and 40;
- data-contract and design validation.
The pretrend diagnostic is interpreted as a falsification test.
Failure to reject pre-treatment coefficients is not recorded as proof that parallel trends holds after treatment.
Dynamic effects¶
The event-study plot asks whether the pricing-policy response changes over time.
In this controlled benchmark the true treatment effect is immediate and constant after adoption.
The appropriate result is therefore a jump around event time zero followed by a roughly flat post-treatment profile.
The case study does not manufacture a lagged effect where none exists.
The repository's separate dynamic-treatment module covers delayed and persistent effects when the DGP contains them.
Heterogeneity¶
For each policy region the case study computes an exploratory region-specific DiD contrast against the mean change in comparison regions.
Those estimates vary because each region is noisy.
The benchmark's true effect is constant, so dispersion in the regional estimates is not evidence of genuine CATE heterogeneity.
A scatter plot against market pressure makes this point visible.
When heterogeneous effects are substantively expected, the repository's causal ML module provides held-out CATE benchmarking, overlap checks, and subgroup stability analysis.
Decision impact¶
The causal effect is translated into:
- incremental orders per treated region-week;
- relative lift versus the post-period comparison mean;
- incremental orders per week across all treated regions;
- cumulative incremental orders over the observed post period;
- an optional illustrative economic value per incremental order.
The value-per-order conversion is a decision input, not part of causal identification.
Changing it changes the economic interpretation, not the estimated treatment effect.
Observational versus experimental route¶
This case study assumes the policy has already been rolled out non-randomly, so Difference-in-Differences is the appropriate primary design.
If the rollout were still under operational control, a stronger route would be to randomize regions or matched regional clusters before deployment.
The repository's experimental-design module supports:
- individual and cluster randomization;
- power and MDE calculations;
- CUPED;
- non-compliance;
- incrementality and lift.
The design should be chosen before the estimator whenever possible.
Reproducible outputs¶
Run:
causal-econometrics-case-study \
--output outputs/retail_pricing_case_study
or:
python examples/case_studies/generate_retail_pricing_case_study.py
The artifact bundle contains:
summary.json
predictive_baseline.json
manifest.json
tables/
decision_impact.csv
estimate_comparison.csv
event_study.csv
placebo_backtest.csv
regional_effects.csv
figures/
outcome_trends.png
event_study.png
regional_effects.png
The figures are generated from code rather than edited manually.
Main figures¶
Regional trends¶
The first figure shows mean orders in policy and comparison regions and marks the intervention week.
It is descriptive: visible separation after treatment is not by itself a causal argument.
Event study¶
The event-study figure shows coefficient estimates and 95 percent intervals by time relative to policy adoption.
Pre-treatment coefficients probe the parallel-trends story. Post-treatment coefficients describe the dynamic effect profile.
Regional-effect diagnostic¶
The third figure shows exploratory region-specific DiD estimates against market pressure.
It is intentionally labelled as a heterogeneity diagnostic rather than a CATE estimate.
Limitations¶
The benchmark is synthetic.
That is useful because the counterfactual truth is known, but it removes many real implementation problems:
- policy timing measurement error;
- spillovers between neighboring regions;
- concurrent commercial interventions;
- changing regional composition;
- missing or revised transaction data;
- policy anticipation;
- endogenous policy timing not captured by fixed regional differences.
A real deployment should preserve the same workflow structure while replacing the benchmark with governed business data and revisiting the identification assumptions from the beginning.