Bayesian and probabilistic causal modelling¶
Bayesian inference changes how uncertainty is represented. It does not replace causal identification.
This module deliberately separates those two ideas.
Randomization, exchangeability, or a stable counterfactual design supplies the causal argument. The Bayesian model supplies a posterior distribution conditional on that design, likelihood, and prior.
Conjugate Bayesian treatment-effect regression¶
For a randomized experiment the model is
[ Y = X\beta + \varepsilon, \qquad \varepsilon \sim N(0,\sigma^2 I). ]
The coefficient vector has a Gaussian prior conditional on sigma squared, and sigma squared has an inverse-gamma prior.
This gives closed-form posterior parameters and direct posterior simulation.
fit_bayesian_treatment_regression returns draws for every coefficient and the residual variance. The treatment coefficient posterior can then be summarized without treating a confidence interval as if it were a probability statement.
Posterior decision quantities¶
summarize_posterior_effect reports:
- posterior mean;
- posterior standard deviation;
- 95 percent credible interval;
- posterior probability that the effect is positive;
- posterior probability that the effect exceeds a user-specified decision threshold.
For example, a statement such as
P(treatment effect > 2 | data, model, prior) = 0.97
is a posterior probability under the model.
That is conceptually different from a frequentist 95 percent confidence interval.
Posterior predictive checks¶
posterior_predictive_check simulates replicated datasets from posterior draws.
It compares:
- observed and replicated means;
- observed and replicated standard deviations;
- posterior predictive tail probabilities;
- RMSE of the posterior-mean prediction.
These checks ask whether the fitted probabilistic model can reproduce basic features of the data it claims to describe.
They are model diagnostics, not causal-identification tests.
Bayesian versus frequentist uncertainty¶
compare_bayesian_frequentist_treatment_effect places two summaries side by side:
- a weak-prior Bayesian credible interval;
- a heteroskedasticity-robust OLS confidence interval.
In a large randomized sample they should usually tell a similar numerical story.
Their interpretations remain different.
The Bayesian interval is a posterior probability statement conditional on the prior and likelihood. The frequentist interval is a repeated-sampling procedure with nominal coverage under its assumptions.
Prior sensitivity¶
prior_sensitivity refits the same treatment-effect model under different prior scales.
This makes prior influence visible.
With little data or a very concentrated prior, the posterior can move materially toward the prior mean. With a large randomized sample and a weak prior, the likelihood dominates.
Prior sensitivity is therefore part of the model report, not an optional afterthought.
Hierarchical regional effects¶
Region-specific randomized effects are noisy when each region has limited sample size.
The hierarchy is
[ \hat\tau_r \mid \tau_r \sim N(\tau_r, s_r^2), ]
[ \tau_r \mid \mu, \omega^2 \sim N(\mu, \omega^2). ]
The model uses the within-region randomized difference in means and its sampling variance as the first level.
A Gibbs sampler then learns:
- each regional effect tau_r;
- the population mean mu;
- the between-region heterogeneity omega.
Noisy region estimates are partially pooled toward the population distribution.
This is not arbitrary shrinkage. The degree of pooling is learned from the relative size of within-region uncertainty and between-region heterogeneity.
The result reports simple split-chain R-hat diagnostics for the population mean and heterogeneity scale.
Why partial pooling matters¶
Unpooled regional estimates can overreact to noise.
Complete pooling hides real heterogeneity.
Partial pooling sits between those extremes:
- regions with precise data move less;
- regions with noisy estimates shrink more;
- the population distribution is estimated jointly.
The simulation contains exact regional treatment-effect truth, so the repository can compare unpooled and pooled estimation error directly.
Bayesian counterfactual time series¶
fit_bayesian_counterfactual uses donor-unit outcomes to model the treated unit during the pre-intervention period.
The donor relationship is estimated with Bayesian linear regression.
Posterior coefficient draws and residual variance generate a posterior predictive distribution for the untreated post-intervention trajectory.
For each posterior draw the intervention effect is
[ \Delta = \frac{1}{T_{post}} \sum_t (Y_t - Y_t^{cf}). ]
The result therefore includes:
- posterior counterfactual mean by period;
- 95 percent predictive bands;
- posterior draws of the average intervention effect;
- posterior probability that the average effect is positive.
As with synthetic control, this requires the pre-treatment donor relationship to remain stable after intervention.
Reproducibility¶
All posterior simulation functions accept explicit random seeds.
The conjugate regression is direct simulation rather than Markov-chain Monte Carlo.
The hierarchical model uses a small Gibbs sampler and reports convergence diagnostics for its population-level parameters.
This keeps the probabilistic machinery inspectable and avoids adding a large probabilistic-programming dependency when the models have tractable conditionals.
Limits¶
Bayesian modelling does not make observational associations causal.
A precise posterior around a biased estimand is still a precise answer to the wrong causal question.
The workflow remains:
- define the estimand;
- justify identification;
- specify a likelihood and prior;
- inspect posterior diagnostics;
- examine prior sensitivity;
- translate the posterior into decision quantities.