Experimental design, incrementality, and lift¶
Randomized experiments identify causal effects by design rather than by statistical adjustment for treatment selection.
The core contrast in this module is between randomized assignment and a targeted observational rollout generated on the same underlying population.
Randomized A/B experiments¶
For assignment A in {0,1}, the intent-to-treat estimand is
[ ITT = E[Y\mid A=1] - E[Y\mid A=0]. ]
Under random assignment, consistency, and no interference, this difference identifies the causal effect of being assigned to treatment.
When treatment compliance is perfect, ITT equals the treatment ATE in the baseline simulation.
When compliance is imperfect, ITT remains the causal effect of assignment and need not equal the effect of treatment receipt.
Stratified randomization¶
Stratified assignment balances treatment within levels of a pre-treatment variable.
The estimator computes treatment-control differences within strata and weights them by stratum sample share.
This can improve precision when the stratification variable is predictive of the outcome.
Cluster randomization¶
Cluster-randomized experiments assign treatment at group level.
The implementation analyzes cluster means as the independent units. This avoids pretending that thousands of individuals are independently randomized when the true number of randomized units is much smaller.
For a causal cluster-level interpretation, spillovers across randomized clusters must be absent or otherwise incorporated into the treatment definition. The generic no-interference contract used by the package is deliberately stronger than this minimum cluster-level requirement.
The power impact of clustering can be approximated with a design effect:
[ DE = 1 + (m-1)\rho, ]
where m is average cluster size and rho is the intraclass correlation.
The power utilities accept a generic design_effect argument so this inflation can be supplied explicitly.
Power, MDE, and sample size¶
For a continuous outcome with equal-sized arms, the normal approximation uses
[ SE(\bar Y_1-\bar Y_0) = \sigma\sqrt{\frac{2DE}{n}}, ]
where n is the sample size per arm.
The repository provides:
- required_sample_size_per_arm;
- minimum_detectable_effect;
- approximate_power.
These are planning approximations, not substitutes for simulation when outcomes, clustering, attrition, or analysis models are complex.
CUPED¶
CUPED uses a pre-treatment covariate X that predicts the outcome:
[ Y_i^{adj} = Y_i - \theta(X_i-\bar X). ]
The coefficient theta is estimated from the covariance between outcome and the pre-period variable.
Because X is measured before treatment, it can reduce residual variance without removing a real treatment effect.
The module reports both the adjusted ITT and the empirical fraction of outcome variance removed by the adjustment.
Incremental lift¶
An experimental treatment effect can be translated into business-facing quantities without changing the causal estimand.
For control mean mu_0 and estimated effect tau:
[ \text{relative lift} = \frac{\tau}{\mu_0}. ]
If one outcome unit is worth r units of revenue, then incremental revenue per eligible unit is
[ r\tau. ]
For an eligible rollout population N, projected total incremental revenue is
[ Nr\tau. ]
These are decision transformations of an experimental effect. They are not a new identification strategy.
Non-compliance¶
The DGP allows assigned users to refuse treatment and control users to cross over.
Then the randomized ITT becomes smaller than the structural treatment effect.
Assignment can be used as an instrument for treatment received. Under relevance, exclusion, monotonicity, and random assignment, 2SLS identifies the complier-average causal effect.
The test suite deliberately shows both quantities side by side:
- ITT answers the effect of assigning the policy;
- the complier effect answers the effect of treatment among units whose treatment status is changed by assignment.
They should not be reported as if they were the same estimand.
Attrition¶
Randomization does not automatically protect a complete-case analysis from post-randomization missingness.
The stress-test DGP makes response probability depend on assignment and on the untreated outcome level.
The resulting complete-case difference in means is biased even though original assignment was randomized.
This is why response rates, missingness mechanisms, and sensitivity analyses belong in experiment review.
Observational comparison¶
The same latent population is also exposed to a targeted non-randomized rollout where treatment probability depends on the pre-period outcome proxy.
A naive treated-versus-control comparison is therefore confounded and differs substantially from the randomized estimate.
This comparison is intentionally direct:
- the experiment identifies an assignment effect through randomization;
- the observational difference is an association until additional identifying assumptions and adjustment methods are supplied.
Experimental methodology contract¶
The complete-case ITT implementation explicitly records:
- consistency;
- random assignment;
- non-informative attrition for the analyzed sample;
- no interference.
For non-compliance, assignment-as-instrument additionally requires relevance, exclusion, and monotonicity for the complier interpretation.