Skip to content

Production causal workflows

The estimator modules in this repository are research components.

A production workflow has additional responsibilities:

  1. validate the data contract;
  2. map business columns into an estimator-ready representation;
  3. enforce the identification design encoded by configuration;
  4. fit the causal estimator;
  5. run assumption diagnostics and temporal backtests;
  6. persist machine-readable results;
  7. record deterministic provenance.

The production package implements that boundary for a common-adoption regional policy workflow.

Retail pricing benchmark

build_retail_pricing_benchmark creates a deterministic business-style dataset with:

  • 48 retail regions;
  • 84 weekly observations;
  • weekly order volume;
  • policy and comparison regions;
  • a common pricing-policy activation week;
  • market-pressure variation.

The benchmark is synthetic by design.

Its untreated potential outcomes and exact treatment effects live in a separate truth table and are never supplied to the workflow.

This makes it possible to validate the entire production pipeline against known truth without claiming that an arbitrary public dataset has a known causal treatment mechanism.

Typed configuration

CommonAdoptionWorkflowConfig contains the complete runtime contract:

  • business column mappings;
  • intervention period;
  • event-study window;
  • placebo backtest dates;
  • minimum pre/post support;
  • artifact options.

The configuration can be loaded from JSON and has a deterministic SHA-256 fingerprint.

No estimator choice depends on notebook state.

An example is stored in:

examples/configs/retail_pricing_policy.json

Schema validation

The lightweight dataframe schema validates:

  • required columns;
  • numeric / integer / binary types;
  • nullability;
  • numeric ranges;
  • composite-key uniqueness.

The common-adoption workflow then adds design-specific checks:

  • treatment-group membership must be constant within region;
  • both policy and comparison regions must exist;
  • the observed treatment path must match the configured common-adoption rule;
  • unbalanced panel support is surfaced as a warning.

Error-level violations raise DataValidationError with the complete machine-readable ValidationReport attached.

Deterministic preprocessing

preprocess_common_adoption_panel maps business columns to the estimator schema:

  • region_id;
  • period;
  • outcome;
  • treated;
  • policy_assigned;
  • event_time.

Rows are sorted with a stable algorithm.

No random imputation, implicit sampling, or notebook state is used.

Estimation and diagnostics

The workflow fits a two-way fixed-effects DiD estimate and an event study.

It then records diagnostics including:

  • number of policy regions;
  • number of comparison regions;
  • pre-treatment period support;
  • post-treatment period support;
  • joint event-study pretrend diagnostic;
  • rejected placebo-date backtests;
  • schema/design warnings.

The pretrend diagnostic is deliberately named pretrend_not_rejected.

A p-value above 0.05 is not stored as "parallel trends passed."

Temporal validation

placebo dates before the true intervention act as temporal backtests.

The same causal design is repeatedly evaluated at dates where no treatment effect can yet exist.

This provides a repeatable check against pre-treatment divergence without inventing a forecasting metric that is unrelated to the DiD estimand.

Provenance

Each workflow run records:

  • canonical input-data SHA-256;
  • configuration SHA-256;
  • workflow version;
  • package version;
  • Python, pandas, and NumPy versions.

The run identifier is derived deterministically from the input hash, configuration hash, workflow version, and package version.

The same inputs and configuration therefore generate the same run ID.

No wall-clock timestamp is used in the reproducibility identity.

Reproducible artifacts

write_workflow_artifacts produces:

config.json
diagnostics.json
event_study.csv
placebo_backtest.csv
summary.json
manifest.json

The manifest records each artifact's SHA-256 and byte size.

The preprocessed panel can optionally be exported as well.

JSON is key-sorted and CSV floating-point serialization is explicit, so artifact content is deterministic for a fixed software environment.

Command-line execution

The Poetry script can execute the complete workflow without a notebook:

causal-econometrics-workflow \
  --input retail_policy.csv \
  --config examples/configs/retail_pricing_policy.json \
  --output outputs/retail_policy

The command validates input, estimates the effect, runs diagnostics and placebo backtests, and writes the result bundle.

Research to production

A research estimator becomes a reusable capability when the methodological code is separated from the operational contract.

In this repository:

  • estimators own statistical calculations;
  • production schemas own input validity;
  • workflow configuration owns the design contract;
  • preprocessing owns deterministic column transformation;
  • diagnostics expose assumption checks programmatically;
  • artifact writers own persistence and provenance.

This separation also makes failures explicit.

A duplicated region-week row should fail before estimation. A policy path that violates the configured design should fail before estimation. A pretrend or placebo warning should remain visible in the artifact bundle rather than being lost in a notebook cell.

Limits

Production engineering cannot rescue a weak causal design.

A perfectly versioned, schema-validated, reproducible workflow can still estimate the wrong causal quantity if the identifying assumptions are implausible.

The production layer therefore preserves methodology diagnostics instead of treating successful software execution as evidence of causal validity.