Regression Discontinuity¶
Regression Discontinuity Design (RDD) uses a treatment-assignment rule that changes discontinuously at a known threshold in a continuous running variable.
The estimand is local: the causal effect for units at, or arbitrarily close to, the cutoff.
Sharp RDD¶
In a sharp design,
[ D_i = 1{R_i \ge c}, ]
where R is the running variable and c is the cutoff.
The local-linear specification used here is
[ Y_i = \alpha + \beta(R_i-c) + \tau 1{R_i\ge c} + \gamma(R_i-c)1{R_i\ge c} + \varepsilon_i. ]
The coefficient tau is the estimated discontinuity at the threshold.
This is not a population ATE. It is a local effect at c.
Fuzzy RDD¶
In a fuzzy design, crossing the threshold changes the probability of treatment but does not determine treatment perfectly.
Threshold eligibility becomes a local instrument:
[ Z_i = 1{R_i\ge c}. ]
The repository estimates fuzzy RDD with local 2SLS, using Z as the instrument for actual treatment D.
Under continuity, local relevance, exclusion, monotonicity, and no precise manipulation, the estimand is a local complier effect around the cutoff.
Bandwidth¶
RDD is inherently local.
A narrow bandwidth reduces dependence on global functional-form assumptions but uses fewer observations.
A wide bandwidth improves precision but can bias a low-order local model if the underlying regression function is curved.
bandwidth_sensitivity estimates the effect over a sequence of windows so the tradeoff is visible rather than hidden.
Running-variable manipulation¶
If units can precisely sort around the threshold, observations immediately above and below the cutoff may no longer be comparable.
density_discontinuity_diagnostic compares local density on the two sides of the cutoff.
It is intentionally lightweight. It is not a full McCrary local-polynomial density test.
The simulation can deliberately move observations from just below to just above the threshold, producing a strong density discontinuity.
Covariate continuity¶
Predetermined covariates should generally evolve smoothly through the cutoff.
covariate_continuity_diagnostic fits the same local-linear discontinuity model to a baseline covariate.
A detected jump is evidence that the two sides differ in ways not plausibly caused by treatment.
Again, failure to detect a jump does not prove the design is valid.
Placebo cutoffs¶
placebo_cutoff_estimates re-estimates a sharp RDD at false thresholds away from the true treatment rule.
Large placebo discontinuities suggest that the estimated treatment jump may be capturing unrelated nonlinear structure or specification error.
Functional-form failure¶
The test suite includes a smooth cubic untreated regression function.
A narrow local-linear fit remains close to the true effect, while a very wide bandwidth uses a poor linear approximation and moves farther from the known threshold effect.
This illustrates why polynomial complexity is not the first remedy for RDD misspecification. Locality and bandwidth choice matter directly.
Identification assumptions¶
Continuity at the cutoff¶
In the absence of treatment, expected potential outcomes must be continuous at the threshold.
Formally,
[ \lim_{r\uparrow c}E[Y(0)\mid R=r] = \lim_{r\downarrow c}E[Y(0)\mid R=r]. ]
The same continuity logic applies to Y(1).
No precise manipulation¶
Units should not be able to sort exactly around the threshold based on latent determinants of the outcome.
Local relevance for fuzzy RDD¶
Crossing the threshold must change treatment probability:
[ \lim_{r\downarrow c}P(D=1\mid R=r) \neq \lim_{r\uparrow c}P(D=1\mid R=r). ]
Exclusion and monotonicity¶
For fuzzy RDD, threshold eligibility is interpreted as an instrument. It therefore also requires the usual local IV logic: eligibility should affect the outcome through treatment, and crossing the threshold should not make some units systematically less likely to take treatment when others become more likely.
Diagnostics are not identification¶
A smooth density, continuous baseline covariates, and stable bandwidth results all strengthen an RDD argument.
None proves continuity of unobserved potential outcomes.
The design remains credible because of the assignment mechanism and the substantive plausibility of local comparability around the cutoff.
Visual diagnostic data¶
rdd_binned_data produces plot-ready local summaries on each side of the cutoff: mean running-variable value, mean outcome, bin size, and side.
The helper deliberately returns data rather than imposing a plotting library or visual style. A case study can layer these points with the fitted local regression lines and the cutoff marker.