Skip to content

Reproducibility

Reproducibility is a package-level design constraint.

Randomness

Scientific code uses independent NumPy generators rather than global random state.

from causal_econometrics.random import make_rng

rng = make_rng(2026)

Simulators accept either a seed or an explicit numpy.random.Generator.

Deterministic workflow identity

Production workflows hash:

  • canonicalized input data;
  • serialized workflow configuration;
  • workflow version;
  • package version.

The resulting run ID does not depend on wall-clock time.

Artifact manifests

Production and case-study workflows write SHA-256 checksums for generated tables, JSON files, and figures.

This lets downstream review distinguish a genuinely reproduced artifact from one that was regenerated with a different configuration or input.

Known truth

For simulation studies, latent truth is kept separate from observed data. Tests and validation code may compare estimates with truth; estimator functions may not use it.

CI

Merge-request pipelines run:

ruff check
ruff format --check
mypy
pytest
mkdocs build --strict

The default-branch pipeline additionally publishes the already-built documentation artifact through GitLab Pages without reinstalling the project.

Reproduction versus identification

Reproducing the same estimate is not evidence that the estimate is causal.

The repository treats these as different requirements:

  • reproducibility: can the same input and configuration regenerate the same analysis?
  • identification: do the design assumptions permit a causal interpretation?

A production workflow should satisfy both, not confuse one with the other.