Skip to content

Causal Econometrics

Identification before estimation.

This repository is a reproducible portfolio of econometric and causal-inference methods for panel, time-series, experimental, and observational data. It is built to expose the entire reasoning chain:

[ \text{question} \rightarrow \text{estimand} \rightarrow \text{identification} \rightarrow \text{estimator} \rightarrow \text{diagnostics} \rightarrow \text{decision} ]

The codebase is deliberately more than a collection of notebooks. Statistical components are typed and tested, simulations expose known causal truth, diagnostics are first-class outputs, and production workflows validate inputs before estimation.

The portfolio-facing case study asks:

What is the incremental effect of a regional pricing policy on weekly customer orders?

The benchmark deliberately creates selective policy assignment. The raw post-policy group difference is therefore badly confounded, while Difference-in-Differences recovers the known causal effect.

Read the retail pricing case study

What is implemented

Capability Examples
Panel econometrics pooled OLS, fixed effects, random effects, clustered uncertainty
Quasi-experiments DiD, event studies, RDD, synthetic control, interrupted time series
Endogeneity IV, 2SLS, weak/invalid-instrument experiments
Observed confounding propensity scores, matching, IPW, outcome regression, AIPW
Experiments A/B testing, CUPED, power/MDE, cluster randomization, incrementality
Heterogeneous effects S/T/X learners, cross-fitted DR forest, CATE validation
Marketing MMM, adstock, saturation, ROI, budget counterfactuals
Bayesian modeling posterior effects, partial pooling, posterior predictive checks
Dynamics distributed lags, carry-over, cumulative treatment effects
Robustness placebos, negative controls, specification curves, hidden-confounding sensitivity
Production schemas, provenance, CLI workflows, deterministic artifacts

Design philosophy

What the repository does not do

It does not treat a treatment coefficient as causal merely because a model returned one. Each causal module declares an estimand and an explicit methodology contract.

Where possible, synthetic data-generating processes make the counterfactual truth known. That lets the repository test not only whether an estimator runs, but when it succeeds and how it fails.

Start here

  1. Install and run the project.
  2. Read Estimands and identification.
  3. Explore the retail pricing case study.
  4. Review the architecture and production workflow.
  5. Use the API overview to locate implementation modules.