Skip to content

Selection on observed covariates

This module studies treatment assignment that is confounded, but where the variables required for exchangeability are observed.

The core identification condition is

[ (Y(1),Y(0)) \perp D \mid X. ]

That condition is stronger than having many covariates. It requires the chosen adjustment set to block all treatment-outcome back-door paths.

Dedicated observed-confounding DGP

The simulator generates two observed covariates, x1 and x2, plus the observable nonlinear transform x1_sq.

Treatment follows a nonlinear logistic assignment model:

[ \operatorname{logit} P(D=1\mid X) = \alpha + \beta_1 x_1 + \beta_2 x_2 + \beta_3 x_1^2. ]

The untreated outcome also depends on x1_sq:

[ Y(0) = \gamma_0 + \gamma_1 x_1 + \gamma_2 x_2 + \gamma_3 x_1^2 + \varepsilon. ]

Treatment adds a constant known effect.

Because x1_sq is available to the analyst, omitting it represents nuisance model misspecification rather than hidden confounding. This gives the repository a clean way to test double robustness.

Propensity scores

The propensity score is

[ e(X)=P(D=1\mid X). ]

fit_propensity_scores estimates it with logistic regression.

A fitted propensity score is not itself evidence that exchangeability holds. It is a balancing device under the assumption that the covariate set is already causally sufficient.

Positivity and overlap

Positivity requires both treatment states to remain possible over the covariate support relevant to the target population.

diagnose_overlap reports:

  • propensity ranges in treated and control groups;
  • their empirical overlap interval;
  • fractions below and above configurable extreme-score thresholds;
  • maximum implied inverse-probability weight;
  • effective sample sizes under weighting.

The assignment_scale parameter can deliberately create near-deterministic treatment assignment. In that setting, estimates may depend heavily on a small number of observations even when the treatment model is correctly specified.

Propensity-score matching

The repository implements one-to-one nearest-neighbour propensity matching with replacement.

Each treated unit is paired to the closest control subject to an optional caliper.

The target is ATT:

[ E[Y(1)-Y(0)\mid D=1]. ]

Matching quality depends on overlap. A close propensity-score match also does not guarantee balance on every covariate, so later production diagnostics should inspect balance explicitly.

The reported paired-difference standard error is a simple descriptive uncertainty summary. Matching with replacement and estimated propensity scores can require more specialised inference in applied work.

Inverse-probability weighting

For ATE, treated observations receive weight proportional to 1/e(X), while controls receive weight proportional to 1/(1-e(X)).

The implementation uses normalised Hájek-style group means:

[ \hat\tau = \frac{\sum_i D_iY_i/e_i}{\sum_iD_i/e_i} - \frac{\sum_i(1-D_i)Y_i/(1-e_i)} {\sum_i(1-D_i)/(1-e_i)}. ]

Propensities are clipped by a configurable numerical trim value. This improves finite-sample stability but also changes the practical target when positivity is poor. Trimming is therefore not a substitute for diagnosing overlap.

Outcome regression

Outcome regression models

[ E[Y\mid D,X]. ]

The current baseline assumes a constant treatment effect. Under a correctly specified outcome model, the treatment coefficient targets ATE in the simulation.

Its failure mode is straightforward: if important nonlinear outcome structure that is also related to treatment is omitted, the treatment coefficient can remain confounded.

Augmented IPW

AIPW combines propensity weighting with separate treated and control outcome models:

[ \psi_i = \mu_1(X_i)-\mu_0(X_i) + \frac{D_i}{e(X_i)}(Y_i-\mu_1(X_i)) - \frac{1-D_i}{1-e(X_i)}(Y_i-\mu_0(X_i)). ]

The estimate is the sample mean of this score.

Under the causal identification assumptions and regularity conditions, AIPW is doubly robust:

  • a correct propensity model can compensate for a misspecified outcome model;
  • a correct outcome model can compensate for a misspecified propensity model;
  • if both nuisance models are wrong, double robustness provides no protection.

The test suite demonstrates all three cases against known truth.

Monte Carlo evaluation

One simulation can be unusually favourable or unfavourable.

evaluate_adjustment_monte_carlo repeats the full DGP and estimation workflow and reports, for each estimator:

  • mean estimate;
  • bias;
  • empirical variance;
  • mean reported standard error;
  • 95 percent confidence-interval coverage.

This makes estimator validation a repeated-sampling exercise rather than a single-example screenshot.

What remains untestable

Good balance, good overlap, and accurate nuisance models do not prove conditional exchangeability.

If an important common cause of treatment and outcome is unobserved, matching, IPW, outcome regression, and AIPW can all be causally biased.

That distinction is why the repository treats diagnostics as evidence about specific assumptions rather than proof of identification.