Skip to content

Sensitivity, falsification, and specification analysis

A causal estimate should survive attempts to make it fail.

This phase collects diagnostics that probe different parts of a causal argument:

  • placebo outcomes;
  • negative-control exposures;
  • placebo intervention dates;
  • sensitivity to omitted confounding;
  • specification curves;
  • compact robustness summaries.

None of these checks proves identification.

Their value is that each can reveal a specific way in which the causal story is inconsistent with the data.

Falsification regression

run_falsification_regression estimates

[ Y^{NC} = \alpha + \tau D^{NC} + X'\beta + \varepsilon. ]

Depending on the design:

  • the outcome can be a placebo or negative-control outcome that treatment should not affect;
  • the exposure can be a negative-control exposure that should not affect the true outcome.

The expected causal effect is zero.

A rejected null is therefore evidence that the falsification check failed.

A non-rejected null is weaker. It can arise because the design is valid, but also because the diagnostic is underpowered.

The repository deliberately records the result as rejects_null rather than "passes identification."

Placebo treatment dates

For Difference-in-Differences, placebo_date_grid moves the intervention date into the true pre-treatment period.

If the treated and control groups were already separating before treatment, false intervention dates can produce non-zero effects.

This is evidence against the untreated-trend story.

It is not proof of post-treatment parallel trends when the placebo dates are null.

Negative controls

Negative controls are most useful when their causal role is justified before looking at the result.

Examples include:

  • a pre-treatment outcome that cannot logically be affected by future treatment;
  • an exposure with similar measurement or confounding structure but no causal pathway to the outcome.

A poorly chosen negative control can be uninformative even if its coefficient is exactly zero.

Partial-correlation sensitivity to unobserved confounding

Observed-covariate adjustment depends on conditional exchangeability.

That assumption cannot be verified from the observed data alone.

confounding_sensitivity_grid makes the hidden-confounding question explicit.

First residualize treatment and outcome on the observed adjustment set:

[ D^ = D - E[D\mid X], \qquad Y^ = Y - E[Y\mid X]. ]

Let a standardized omitted confounder U have residual correlations

[ \rho_D = Corr(D^,U), \qquad \rho_Y = Corr(Y^,U). ]

The coefficient of D after including U can be written as

[ \beta_D(U) = \frac{sd(Y^)}{sd(D^)} \frac{ r_{YD}-\rho_Y\rho_D }{ 1-\rho_D^2 }. ]

The grid evaluates this coefficient over user-specified combinations of rho_D and rho_Y.

The three pairwise correlations must define a valid correlation matrix. Impossible combinations are retained in the output but marked invalid rather than silently interpreted.

Symmetric correlation needed to null

For a simple symmetric benchmark in which the omitted confounder has equal correlation magnitude with residualized treatment and outcome, the association can be reduced to zero when approximately

[ |\rho_D| = |\rho_Y| = \sqrt{|r_{YD}|}, ]

with the correlation signs chosen to explain the observed association.

The returned symmetric_correlation_to_null is therefore a scale for discussion, not a universal robustness value.

It does not replace domain knowledge about whether an omitted variable of that strength is plausible.

Specification curves

A specification curve should not be a random search over models until one produces a desired result.

run_adjustment_specification_curve accepts a pre-declared list of AdjustmentSpecification objects.

Each specification records:

  • estimator family;
  • propensity-score adjustment set where relevant;
  • outcome-regression adjustment set where relevant.

The current curve supports:

  • outcome regression;
  • IPW;
  • AIPW.

The output retains the modelling choices beside the estimate, uncertainty, and zero-exclusion indicator.

This makes disagreement across specifications inspectable rather than hiding it behind one selected model.

Defensible specifications

A large set of specifications is not automatically robust.

A useful specification set should differ only along modelling choices that are scientifically defensible.

Examples include:

  • alternative pre-specified covariate sets;
  • plausible nonlinear terms;
  • different valid adjustment estimators;
  • reasonable trimming or overlap rules.

Known invalid adjustment sets should be shown as failure cases, not counted as evidence of robustness.

Robustness summary

build_robustness_summary converts a specification curve and falsification results into a compact machine-readable report:

  • number of specifications;
  • median, minimum, and maximum effect estimate;
  • fraction of estimates with the same positive sign;
  • fraction whose confidence interval excludes zero;
  • number of falsification checks;
  • number of rejected falsification nulls.

This is designed for dashboards or stakeholder reporting.

It is not a score and should not be interpreted as a probability that the causal claim is true.

When the effect is not identified

Sensitivity analysis has an important stopping point.

Suppose an unobserved variable can be arbitrarily related to both treatment and potential outcomes.

Without an additional design assumption, instrument, experiment, longitudinal restriction, or external information, the observed treatment-outcome distribution does not point-identify the ATE.

No amount of:

  • propensity-score modelling;
  • flexible machine learning;
  • bootstrap precision;
  • specification searching;
  • Bayesian posterior concentration

can recover identification that is absent from the design.

In that situation the correct scientific conclusion is that the effect is not identified under the available assumptions.

Sensitivity analysis can show what additional confounding strength would change the estimate, but it cannot manufacture the missing counterfactual information.

What each check can establish

Check Can reveal Cannot establish
Placebo outcome treatment predicts an outcome it should not cause absence of all unmeasured confounding
Negative-control exposure residual association inconsistent with the design that the main treatment is exogenous
Placebo date pre-treatment divergence future parallel trends
Confounding sensitivity how strong a stylized omitted confounder must be to change the estimate actual strength of an unobserved variable
Specification curve dependence on defensible modelling choices correctness of every specification
Robustness summary disagreement and failed diagnostics in compact form probability that the causal claim is true