Skip to content

Causal machine learning and heterogeneous treatment effects

Causal machine learning is useful when the scientific question concerns how a treatment effect varies with observed characteristics.

The target is not ordinary prediction.

For covariates X, the conditional average treatment effect is

[ \tau(x) = E[Y(1)-Y(0)\mid X=x]. ]

A model can predict outcomes accurately and still estimate treatment-effect heterogeneity badly.

This module therefore evaluates factual prediction and CATE recovery with separate metrics.

Heterogeneous-effect benchmark DGP

The dedicated simulator exposes four observed covariates and a nonlinear, interaction-rich true CATE:

[ \tau(X) = \tau_0 + a(0.8x_1-0.5x_2) + b\sin(x_3) + c x_1x_2. ]

Treatment assignment also depends on X, so causal learners must operate under observed confounding rather than randomized treatment.

The truth table includes:

  • true propensity;
  • conditional untreated mean;
  • conditional treated mean;
  • exact CATE.

The learner never receives the true CATE during fitting.

S-learner

The S-learner fits one outcome model

[ \hat\mu(X,D) ]

using treatment as another feature.

CATE is obtained by toggling treatment:

[ \hat\tau(X) = \hat\mu(X,1)-\hat\mu(X,0). ]

It is simple and data-efficient, but a flexible predictor can still choose to use treatment weakly when treatment heterogeneity is subtle relative to outcome variation.

T-learner

The T-learner fits separate outcome models:

[ \hat\mu_1(X) \quad\text{and}\quad \hat\mu_0(X). ]

Then

[ \hat\tau(X) = \hat\mu_1(X)-\hat\mu_0(X). ]

It can represent treatment-specific response surfaces naturally, but each model uses only one treatment arm.

X-learner

The X-learner first fits the T-learner outcome models.

For treated units it imputes

[ D_i^1 = Y_i-\hat\mu_0(X_i), ]

and for controls

[ D_i^0 = \hat\mu_1(X_i)-Y_i. ]

Separate models are fitted to these imputed effects and combined using the estimated propensity score.

This can be especially useful when treatment-group sizes are imbalanced.

Cross-fitted DR forest

The repository's forest-based heterogeneous-effect estimator is intentionally not labelled a formal causal forest.

It first constructs cross-fitted doubly robust pseudo-outcomes:

[ \psi_i = \hat\mu_1(X_i)-\hat\mu_0(X_i) + \frac{D_i}{\hat e(X_i)} [Y_i-\hat\mu_1(X_i)] - \frac{1-D_i}{1-\hat e(X_i)} [Y_i-\hat\mu_0(X_i)]. ]

Nuisance predictions are generated out of fold.

A random forest then models

[ E[\psi\mid X]. ]

This creates a transparent forest-based DR learner: orthogonalization is explicit, cross-fitting is explicit, and the forest is used only after the causal score has been constructed.

A dedicated generalized random forest implementation could be added later, but the current method avoids presenting an ordinary random forest as if it were a causal forest.

Identification still comes first

All learners share the same causal identification contract:

  • consistency;
  • conditional exchangeability;
  • positivity;
  • no interference.

Flexible ML does not repair hidden confounding.

If treatment and potential outcomes remain confounded after conditioning on X, a more powerful learner estimates a more flexible biased association.

CATE metrics

When exact treatment-effect truth is available, evaluate_cate reports:

  • CATE RMSE, often called PEHE-style error in simulation work;
  • CATE MAE;
  • Pearson correlation;
  • Spearman rank correlation;
  • calibration intercept;
  • calibration slope.

Ranking and calibration are both important.

A model may rank high-benefit units correctly while systematically exaggerating effect magnitudes, or be well calibrated on average while ranking subgroups poorly.

Prediction metrics are separate

evaluate_factual_prediction reports:

  • factual outcome RMSE;
  • factual outcome MAE;
  • factual outcome R-squared.

These values are deliberately stored in a different metric object.

A low factual RMSE is not evidence of accurate treatment-effect heterogeneity.

benchmark_causal_learners uses a held-out split and reports the two metric families in separate columns.

Overlap-aware evaluation

CATE estimation becomes unstable where treatment is almost deterministic.

evaluate_cate_with_overlap compares treatment-effect performance:

  1. over the full evaluation sample;
  2. only where the estimated propensity lies inside a specified overlap range.

It also reports the fraction of observations retained.

This is a diagnostic, not permission to silently discard difficult regions of the population. Changing the overlap population changes the practical target of the analysis.

Subgroup stability

subgroup_cate_summary sorts observations into quantiles of predicted CATE and reports:

  • predicted mean effect;
  • true mean effect;
  • subgroup bias;
  • subgroup size;
  • bootstrap standard deviation of the true subgroup mean.

The purpose is to ask whether apparently high-value and low-value treatment groups remain distinct, rather than judging heterogeneity from an individual scatter plot alone.

Model choice

Random forests are used here because they can represent nonlinearities and interactions while keeping the meta-learning logic visible.

The statistical lesson is not that forests are intrinsically causal.

The causal content comes from:

  • the estimand;
  • the identification assumptions;
  • the nuisance-model strategy;
  • orthogonalization/cross-fitting where applicable;
  • overlap;
  • held-out treatment-effect validation when truth is available.