This section covers the structured abacus.pipeline runner: how it loads a
config and dataset, executes the retained stage sequence, and writes
reproducible run artefacts to disk.
Pages
Runner Overview - How run_pipeline(...) works, which
stages run, and when the optimisation stage is skipped.
YAML Configuration - Which YAML keys the runner
consumes and how they map to model build, data loading, holidays, and
optimisation.
Blocked Holdout Validation - What Stage 35
does, how to configure it, and how to read the holdout metrics and plots.
CLI Reference - The thin python -m abacus.pipeline.runner
interface and its supported flags.
Output Directory Schema - The run directory
layout, manifest schema, stage statuses, and main artefacts.
Extending the Runner - How to add a stage or wire
in reporting without bypassing the manifest and artifact helpers.
Subsections of Pipeline Runner
Runner Overview
Use the pipeline runner when you want a full disk-backed PanelMMM run instead
of only an in-memory fit.
The runner loads a YAML config and a CSV dataset, builds the model, executes a
fixed stage sequence, and writes each stage’s artefacts into a structured run
directory. When validation is enabled, the runner performs a second train-window
fit for the blocked holdout stage, so the run takes longer than a pure
full-sample fit.
The path to run_manifest.json inside that directory
What the runner does
run_pipeline(...) performs these steps:
Load the YAML config with load_yaml_config(...).
Load X and y from CSV using load_pipeline_data(...).
Merge CLI sampler overrides with YAML fit through
build_model_kwargs(...).
Create the output directory tree and initialise run_manifest.json.
Run the retained stages in order, updating the manifest after every stage.
For named estimators, Stage 00 prepares the model and stores it in the shared
PipelineContext; Stage 10 completes the deferred graph before prior
predictive sampling. Configurations without a named estimator use the direct
builder path, which can complete the graph in Stage 00. A populated
model_class field identifies the model object, not graph completion.
Runner-only roots such as prior_sensitivity, ai_advisor, diagnostics,
and validation stay on the
pipeline context and are stripped before the public MMM builder validates the
model YAML.
Stage order
This is the canonical stage sequence. The runner uses a fixed stage list;
optional stages retain their place and record a skipped status when disabled.
Stage 80 inventories evidence and stage statuses. It does not generate
automated business conclusions or certify decision readiness.
Stage key
Directory
Purpose
Optional
metadata
00_run_metadata
Prepare the model and write resolved config and dataset metadata
No
prior_sensitivity
05_prior_sensitivity
Write resolved prior-sensitivity scenario configs and manifests
Yes
ai_advisor
08_ai_advisor
Write privacy-safe deterministic and optional LLM guidance artifacts
Yes
preflight
10_pre_diagnostics
Complete any deferred graph; draw and plot prior predictions
No
fit
20_model_fit
Fit the model, save InferenceData, write trace and summary
Raw input screening, MCMC, predictive, and residual diagnostics
No
ai_diagnostics_advisor
55_ai_diagnostics_advisor
Write privacy-safe LLM diagnostics review artifacts for enabled AI advisor runs
Yes
curves
60_response_curves
Saturation-only, forward-pass direct contribution, and adstock curve artefacts
No
optimisation
70_optimisation
Budget optimisation artefacts
Yes
interpretation
80_interpretation
Inventory retained evidence for analyst review
No
The prior-sensitivity stage is marked skipped when the YAML config does not
contain prior_sensitivity or it is disabled. The AI advisor stage follows the
same convention for ai_advisor. The diagnostics advisor stage runs by default
for enabled ai_advisor blocks and is marked skipped only when ai_advisor
is absent, disabled, or has diagnostics_review_enabled: false. The validation
stage is marked skipped when the YAML config does not contain validation or
it is disabled. The optimisation stage is also optional; it returns None and
is marked skipped when the YAML config does not contain an optimization
block.
PipelineRunConfig controls runtime settings that sit outside the YAML model
specification.
Field
Purpose
config_path
YAML file to load
output_dir
Root directory under which the run directory is created
run_name
Optional run-name override; otherwise the config filename stem
dataset_path
Optional combined dataset CSV override
x_path, y_path
Optional feature and target CSV overrides
holidays_path
Optional holiday CSV override
target_column
Target column name used during CSV loading
prior_samples
Number of prior predictive samples for Stage 10
draws, tune, chains, cores, random_seed
Sampler overrides merged onto YAML fit
curve_samples, curve_points
Curve sampling settings for Stage 60
The draws, tune, chains and cores overrides do not necessarily bound
Stage 35: explicit validation.sampler values take precedence for the refit.
Use the bounded software smoke
for an execution check with validation explicitly skipped.
Only sampler settings are merged into model construction. Other overrides are
used by the runner itself during data loading, holiday resolution, diagnostics
reporting, and output setup.
The timestamp is generated in UTC. Each invocation exclusively allocates a new
run directory, including simultaneous invocations with the same name and
timestamp. The random suffix is opaque; use PipelineRunResult.run_dir and
manifest_path instead of reconstructing paths from the name and timestamp.
Existing runs are never reused or resumed by this allocator.
A run name must be a non-empty filename component, without path separators.
Set output_dir to choose its parent. Allocated run directories have owner-only
permissions on POSIX (0700); manifest files have mode 0600.
All stage directories are created up front, even if a later stage is skipped or
the run aborts. An allocation or stage-directory setup error propagates; a
partially initialised new run directory may remain for inspection.
Manifest updates are published by replacing the previous file with a completed
temporary file in the same directory. Readers see a complete old or new
snapshot. A publication failure propagates and leaves the previous snapshot
intact, which may therefore be stale. This does not make stage artefacts
transactional or provide power-loss durability.
Failure and skip behaviour
If a stage raises an exception:
the current stage is marked failed
the run manifest is marked failed
all still-pending later stages are marked not_reached
run_pipeline(...) re-raises the exception
If a stage returns None:
the stage is marked skipped
the manifest warning records that no configuration was supplied for that
optional stage
Reporter hook
run_pipeline(...) accepts an optional reporter that implements the
PipelineReporter protocol.
The pipeline runner reads the same YAML model specification used by
build_mmm_from_yaml(...), then adds a small set of runner-specific conventions
for data loading, prior-sensitivity planning, optional AI advisor guidance,
optional blocked holdout validation, and Stage 70 optimisation.
This page documents the keys that the runner actually consumes.
Root keys
Key
Required
Used for
data
Usually
Resolve dataset paths when you do not pass dataset_path, x_path, or y_path through PipelineRunConfig
target
Yes
Define the target column and business target type
estimator
No
Declare a named estimator preset; time_series, fe, and cre are currently released
dimensions
No
Declare panel-dimension columns such as geo or brand
media
Yes
Define channel/control columns and transform types
scaling
No
Configure target/channel scaling rules
effects
No
Append additive effects in YAML order before build_model(...)
priors
No
Override model-level priors and prefixed transform priors
fit
No
Default sampler settings for Stage 20 fitting
holidays
No
Add holiday events before model build
original_scale_vars
No
Add original-scale contribution variables before fitting
inference_data
No
Attach existing InferenceData when the file exists
prior_sensitivity
No
Write a pre-fit scenario plan for prior robustness checks
ai_advisor
No
Write privacy-safe AI advisor guidance before model fitting
Relative paths in YAML are resolved relative to the YAML file’s directory.
diagnostics is runner-only. The structured pipeline reads it, but
build_mmm_from_yaml(...) still validates only the public MMM model schema.
prior_sensitivity and ai_advisor are also runner-only. They are consumed by
the structured pipeline before model fitting and stripped before the public MMM
YAML builder validates the model specification.
validation is also runner-only. The structured pipeline reads it for Stage 35
blocked holdout scoring, but the public MMM YAML builder never sees it.
Core modeling blocks
The runner always builds a PanelMMM, so the public YAML no longer exposes a
model.class field. Instead, it reads:
data.date_column
target.column
target.type
media.channels
media.controls, if any
estimator, for a named preset
dimensions.panel, if any
media.adstock
media.saturation
fit
estimator
The released named single-series contract is:
estimator:type:time_series
It requires one observation per date and no panel unit. It builds the same
single-series graph as the established configuration with no
dimensions.panel for the same configuration. With the default model settings,
this means one global intercept and shared media, control, adstock, saturation,
and residual parameters. Other explicit single-series model options retain
their established behaviour; the estimator declaration does not silently
override them.
It accepts one unit column and uses an exact within-unit orthonormal-contrast
likelihood. Unit intercepts are absorbed. Media and control slopes, adstock,
saturation, and residual scale are shared across units. The FE preset does not
support common time effects, annual seasonality, custom additive effects, or
time-varying parameters. See Fixed-effects Estimator
for the estimability checks and interpretation limits.
The released correlated-random-effects contract is:
It accepts one balanced unit panel. Media and control slopes, geometric adstock,
logistic saturation, and residual scale are shared across units. The graph uses
an exact marginal Gaussian random-intercept likelihood and adds centred unit
means of the transformed media basis and eligible time-varying controls. Common
time effects, seasonality, custom effects, calibration, optimisation and
fixed-budget scenario optimisation are not supported. Prediction and
historical/manual scenarios require all fitted units and reject unseen units
and unit subsets. Manual CRE scenarios retain the fitted training-period
Mundlak summaries rather than recomputing them from planned spend. See
Correlated-random-effects Estimator
for the estimability and interpretation limits.
The re declaration validates as typed configuration but remains release-gated.
It fails before graph construction and does not fall back to the advanced
panel-dimension surface.
Do not combine estimator with dimensions.panel. Abacus rejects the mixed
declaration rather than guessing which semantics you intended.
data
The runner loads data before building the model. It supports two CSV layouts.
Combined dataset
data:dataset_path:"dataset.csv"
The runner reads the CSV, removes the target column from X, and uses that
column as y.
Separate feature and target files
data:x_path:"X.csv"y_path:"y.csv"
When loading y_path:
if the configured target column exists, the runner uses that column
otherwise, if the file has exactly one column, the runner uses that column
and renames it to the target name
Target column resolution
The runner resolves the target column in this order:
PipelineRunConfig.target_column or CLI --target-column
target.column
"y"
Use the CLI override only when you want to change how the runner reads the CSV.
Keep it consistent with target.column in YAML.
fit
fit controls Stage 20 fitting because the fit stage
calls:
The CLI or PipelineRunConfig.holidays_path overrides holidays.path.
If you omit both path and the override but still configure holidays,
Abacus falls back to the bundled abacus.data:holidays.csv.
Country-selection rules:
time-series configs default to US when holidays.countries is omitted
geo-panel configs must declare holidays.countries explicitly
geo-panel configs must provide multiple countries, for example
["UK", "FR", "DE"]
If you provide a catalogue-style holiday CSV, Abacus only creates holiday
effects for the countries listed in holidays.countries.
Holiday modes:
event creates one latent holiday/event effect per
holiday row, which is why posterior summaries include terms like
holiday_effect_size[...].
pooled_control creates one pooled binary holiday regressor over time and
estimates a single shared holiday coefficient. This is useful when you want
a strict calendar-only single holiday term instead of one parameter per holiday.
prophet_component fits Prophet on the training target with the configured
holiday calendar, extracts the continuous holidays component, and uses that
single smoothed series inside the MMM as one holiday term. For panel models,
Abacus fits one Prophet holiday component per panel series and filters the
holiday calendar by geo when that dimension is present.
For holiday effects, each model date labels the start of its observed period.
Abacus assigns an inclusive holiday date range to every model period it
overlaps. For example, on W-MON data, a Wednesday or Sunday holiday is
assigned to the Monday date that starts that week. Daily data retains its
existing date-by-date assignment.
Default behavior:
configs default to prophet_component
use event explicitly when you want one latent holiday effect per holiday row
Current limitation:
pooled_control currently supports only single-country, non-geo models.
prophet_component requires exactly one holiday country unless the model has
a geo dimension, in which case it can route multiple holiday countries to
the matching geo-level panel series.
original_scale_vars
Use original_scale_vars when you want specific contribution variables to be
available on the original target scale:
original_scale_vars:- channel_contribution- y
The builder applies these through
model.add_original_scale_contribution_variable(...) before fitting.
inference_data
inference_data.path is passed through to the YAML builder. If the file exists,
Abacus attaches that InferenceData when the build completes: in Stage 10 for
named estimators with a deferred graph, or Stage 00 for the direct builder
path. See the runner lifecycle.
Important: the structured runner still executes Stage 20 and fits the model
again. inference_data.path does not currently skip fitting.
prior_sensitivity
Use the optional prior_sensitivity block when you want the runner to write a
pre-fit prior scenario plan. This stage does not fit every scenario. It creates
resolved scenario configs that can be reviewed, approved, and run deliberately.
prior_sensitivity:enabled:truescenario_policy:manualreference:referencescenarios:reference:description:Current approved prior specification.tighter_media_effect:description:Lower media-effect amplitude on the scaled target space.overrides:media.saturation.priors.beta:distribution:HalfNormalsigma:0.5dims:["channel"]
Supported keys:
Key
Meaning
enabled
Set to true to write Stage 05 prior-sensitivity artifacts
scenario_policy
manual for declared scenarios or conservative_mmm for generated relative scenarios
reference
Scenario name for the unchanged reference config
scenarios
Optional manual scenario declarations
allow_model_structure_overrides
Required before scenarios can change transform structure such as media.adstock.l_max
Scenario names are slugs such as reference, longer_memory, or
tighter_media_effect. Avoid names such as baseline; in MMM, baseline has a
model meaning and should not be overloaded as a scenario label.
Allowed override paths are intentionally narrow:
media.adstock.priors.*
media.saturation.priors.*
priors.*
selected transform-structure paths such as media.adstock.l_max, only when
allow_model_structure_overrides: true
The stage writes both a human-readable manifest and an LLM-safe manifest. Use
the LLM-safe file when passing scenario context to an external model because it
aliases override paths and avoids free-text descriptions.
ai_advisor
Use the optional ai_advisor block when you want privacy-safe, evidence-grounded
modelling guidance from deterministic rules and, optionally, OpenAI or
OpenRouter. The advisor proposes controlled tests. It does not approve a model,
establish causal identification, or apply a config patch to the run config.
autopilot; the advisor prioritizes concise recommendations and approval-ready options
privacy
anonymized_relative; raw channel names and raw business values are excluded from the LLM payload
approval
file_based; proposed config changes are written as files for user approval
write_outputs
Set to false to disable artifact writes even when the block is enabled
llm_enabled
Set to false to run deterministic privacy/rule checks without an LLM call
diagnostics_review_enabled
Defaults to true; set to false to skip the post-fit 55_ai_diagnostics_advisor LLM review after structured diagnostics
openai_model
OpenAI model name used for the advisor call
openai_timeout_seconds
Request timeout for the OpenAI call
openrouter_model
OpenRouter model name used for the advisor call
openrouter_timeout_seconds
Request timeout for the OpenRouter call
The pipeline reads OPENAI_API_KEY or OPENROUTER_API_KEY from the process
environment based on provider. For local development, an untracked repo-root
.env file is also supported. Do not commit API keys.
The advisor stages complete even if an LLM call fails. In that case they write
an error artifact and the rest of the pipeline can continue. Deterministic
rules provide a minimum decision state: an LLM may make the state stricter, but
cannot override a failed gate or weak-identification warning with a more
favourable conclusion.
When the advisor proposes a valid config patch, Stage 08 writes:
The approval command writes approved_config.resolved.yaml and
approval_record.yaml beside the advisor artifacts. It does not mutate the
source YAML config.
By default, an enabled ai_advisor block also runs
55_ai_diagnostics_advisor after structured diagnostics. That post-fit advisor
uses anonymized channel aliases, convergence counts, normalized predictive
metrics, coverage metrics, and scale-free design diagnostics. It intentionally
excludes raw target-scale fit errors from the LLM payload. Set
diagnostics_review_enabled: false to run only the pre-fit advisor.
optimization
Add an optimization block when you want Stage 70 to run. If this block is
absent, Stage 70 is marked skipped.
The YAML builder validates this block when the config is loaded. start_date
and end_date are always required, and you must provide exactly one of:
optimization.budget for the preferred user-facing budget spec
optimization.total_budget for the legacy per-period budget input
Preferred user-facing budget spec: absolute or relative
total_budget
None
Legacy per-period budget input kept for backward compatibility
response_variable
total_media_contribution_original_scale
Optimisation objective variable
budget_distribution_over_period
None
Time weights over the optimisation window
budget_bounds
Derived or default
Explicit spend bounds
minimize_kwargs
None
Options forwarded to the existing SciPy minimiser contract; defaults remain SLSQP, ftol=1e-9, maxiter=1000
spend_constraint_lower
0.3 when deriving bounds
Relative lower bound around scaled reference spend
spend_constraint_upper
0.3 when deriving bounds
Relative upper bound around scaled reference spend
default_constraints
true
Whether to add the default equality budget constraint
noise_level
0.0
Must be zero for the deterministic allocation comparison
include_last_observations
false
Must be false, matching the optimiser initial history
include_carryover
true
Must be true, matching the optimiser response horizon
Stage 70 rejects other values for these three comparison options. Both plans
use the same posterior draws, time profile, zero initial history and full
carryover horizon. The current row replays reference-window spend totals
under that profile; it is not historical attribution. The comparison summarises
expected media contribution, without observation noise.
Budget spec modes:
budget.mode: absolutebudget.value is total spend over the full optimisation horizon.
budget.mode: relativebudget.value is a multiplier on the chosen basis.
budget.basis: reference_window_total
Abacus resolves the budget against the same reference-window total spend it
already uses for current-plan comparison and default bound derivation.
Important budget-unit note
The preferred optimization.budget block uses total horizon spend. Stage
70 converts that to the wrapper’s per-period contract internally before calling
PanelBudgetOptimizerWrapper.optimize_budget(...).
The legacy optimization.total_budget field is still supported, but it keeps
the old wrapper-facing per-period spend contract.
Time profiles must include every budget dimension with unique labels matching
the model exactly. Labels and dimensions can be reordered; the optimiser aligns
them before use. Fractions must be finite and non-negative and sum to one along
date for every budget cell. Do not omit labels or use negative weights to
balance a column total.
The same shape works for budget_bounds, but with an additional "bound"
dimension containing "lower" and "upper".
diagnostics
Stage 50 resolves a complete, versioned decision-gate profile and writes it to
50_diagnostics/diagnostic_gates.resolved.yaml. The packaged default is
abacus/pipeline/diagnostic_gates.default.yaml. A documented copy is available
at examples/diagnostic_gates.team.yaml. Copy that file when your team needs a
governed profile with different thresholds; keep the source profile in version
control with the model configuration.
Use gates_file to select that profile. Relative paths are resolved from the
model YAML file. Optional inline thresholds take precedence over the selected
profile and are recorded in the resolved artifact.
equality triggers the relevant warn or fail boundary; the zero-divergence
gate is the explicit exception, where zero passes and any positive count fails
a selected gate file must be schema version 1 and define every supported gate
omit the block entirely to use the packaged default profile
These gates classify available diagnostic evidence. Passing them does not prove
parameter identification, prior robustness, model validity, or causal
identification. In particular, VIF and condition number are raw-design screens.
A warning indicates weak-identification risk and should trigger controlled
reparameterisation or prior-sensitivity runs. A clean screen only means that no
material warning was detected by those checks.
This block affects only the structured runner. It is stripped before Stage 00
model preparation so the public MMM YAML schema remains unchanged.
validation
Use the optional validation block when you want Stage 35 blocked holdout
scoring. This is the runner’s out-of-sample tail check: Abacus refits a clean
model on the earlier dates and scores only the final blocked window.
Set to false to skip Stage 35 while keeping the stage in the manifest
holdout_observations
Number of unique dates to reserve for the blocked holdout window
include_last_observations
Keep lag history for carryover-sensitive holdout scoring
coverage_levels
Coverage levels reported in Phase 10; use the fixed 50, 80, and 94 percent defaults
sampler
Optional validation-only sampler overrides for the train-window refit
Validation settings are merged in this order: YAML fit, runner sampler
overrides, then validation.sampler. The last value wins. Reducing CLI draws
or tuning does not override an explicit validation budget.
Phase 10 reports coverage as coverage_50, coverage_80, and
coverage_94. Keep those defaults unless the implementation and tests are
updated together.
The validation stage builds a clean train-window model for holdout scoring and
ignores inference_data.path so the refit does not inherit attached posterior
state from the main model build.
For a full explanation of why the split is blocked, how to read crps and
coverage, and rules of thumb for weekly MMM, see
Blocked Holdout Validation.
Override precedence
For the runner, precedence is:
Setting
Higher precedence
Lower precedence
Combined dataset path
dataset_path / --dataset-path
data.dataset_path
Split CSV paths
x_path, y_path / --x-path, --y-path
data.x_path, data.y_path
Holiday CSV path
holidays_path / --holidays-path
holidays.path
Sampler settings
PipelineRunConfig or CLI overrides
fit
Target column for CSV loading
target_column / --target-column
target.column, then "y"
Diagnostics thresholds
diagnostics.thresholds
retained Stage 50 defaults
Common pitfalls
Using Parquet paths in the pipeline data block. The runner data loader reads
CSV only.
Providing only one of data.x_path or data.y_path.
Mixing the preferred horizon-based optimization.budget block with the
legacy per-period optimization.total_budget field.
Assuming diagnostics is part of the public MMM builder schema. It is a
runner-only block.
Assuming inference_data.path skips Stage 20 fitting. It does not.
Forgetting that relative paths are resolved from the YAML file directory, not
from the shell working directory.
For large original-currency objectives, record an explicitly justified numerical
tolerance in optimization.minimize_kwargs, for example
{options: {ftol: 0.000001, maxiter: 1000}}. Verify feasibility and objective
accuracy independently; successful termination alone is not statistical
validation. Abacus does not automatically retry failed solvers with looser
tolerances. The underlying Python wrapper accepts the same minimize_kwargs.
Output Directory Schema
Each pipeline run creates a timestamped directory under the configured
output_dir:
The timestamp is generated in UTC. A random suffix and exclusive directory
creation keep same-name, same-time runs separate. Use the returned run path;
do not infer a path from the timestamp. Existing saved runs remain readable.
On POSIX, newly allocated run directories have mode 0700 and manifests have
mode 0600.
The runner creates every stage directory up front, then publishes complete
run_manifest.json snapshots by atomic replacement as stages start, complete,
skip, or fail. Each run has its own manifest and artefacts.
scenario_planner/recipes/ is a post-fit evidence area, not a pipeline stage.
It appears only after a user evaluates a retained scenario recipe. Recipe
evaluation does not mutate run_manifest.json or refit the model.
Stage directories
Stage
Directory
Typical artefacts
metadata
00_run_metadata
resolved config, model metadata, and estimator contract
prior_sensitivity
05_prior_sensitivity
scenario configs, human manifest, and LLM-safe manifest
Each recipe evaluation creates a new directory under
scenario_planner/recipes/. It contains the resolved request, validation
evidence, fitted estimator manifest, allocation tables, posterior
media-contribution summaries with 94% highest-density intervals, a versioned
dashboard payload, and a SHA-256 artefact manifest.
See Comparison Outputs for the
complete file contract and interpretation boundary.
Main artefacts by stage
00_run_metadata
Main files:
a copy of the original config under its source filename
config.original.yaml
config.resolved.yaml
session_info.txt
dataset_metadata.json
model_metadata.json
data_dictionary.csv
design_matrix_manifest.csv
spec_summary.csv
estimator_summary.txt and estimator_manifest.yaml for named estimators
estimator_estimability.csv for the resolved estimator screen summary
holiday_feature_manifest.csv when holidays are configured
config.resolved.yaml normalises configured data and holiday paths to absolute
paths and records the effective sampler configuration on the model.
05_prior_sensitivity
Main files:
scenario_manifest.yaml
llm_safe_scenario_manifest.yaml
<scenario_name>/config.resolved.yaml for each generated or declared scenario
This stage is optional. When prior_sensitivity is absent or disabled in YAML,
the directory still exists and the stage is marked skipped.
scenario_manifest.yaml is the local, human-readable manifest. It can include
scenario descriptions and raw override paths. Use
llm_safe_scenario_manifest.yaml when passing scenario context to an external
LLM because it aliases override paths and avoids free-text scenario prose.
08_ai_advisor
Main files:
evidence.json
rules_summary.json
advisor_status.json
advisor_response.json, when an LLM advisor call succeeds
<provider>_response.raw.json, when an LLM advisor call succeeds
advisor_recommendations.md, when an advisor response is available
config_patch_proposal.yaml, when the advisor proposes a YAML patch
approval_request.yaml, when the proposed patch is valid and ready for file-based approval
advisor_error.json, when the LLM call fails
config_patch_error.json, when an advisor patch is invalid or unsafe
This stage is optional. When ai_advisor is absent or disabled in YAML, the
directory still exists and the stage is marked skipped.
The advisor evidence is privacy-safe by design: it aliases media channels,
prior paths, and prior specs, and excludes raw business values. Deterministic
rule checks always run before the optional LLM call. If those rules mark the
payload as unsafe, the LLM call is blocked and the status artifact records the
reason.
Patch proposals are validated before an approval request is written. Supported
patches use an overrides: mapping with safe prior paths only. After a user
sets approval_request.yaml to status: approved, the file-based approval CLI
writes approved_config.resolved.yaml and approval_record.yaml in this
directory.
10_pre_diagnostics
Main files:
prior_predictive.nc
prior_predictive.png
fixed_effects_estimability.csv and
fixed_effects_estimability.json for FE
cre_structural_estimability.json,
cre_reference_estimability.json, and
cre_reference_estimability_features.csv for CRE
20_model_fit
Main files:
model.nc
trace.png
posterior_summary.csv
cre_postfit_estimability.json for CRE
posterior_summary.csv is intentionally compact. It summarizes structural
posterior parameters such as adstock, saturation, seasonality, holiday, and
likelihood terms, and omits per-date deterministic series like weekly channel
contributions or fitted paths. Use the assessment and decomposition stages for
time-indexed fitted or contribution outputs.
30_model_assessment
Main files:
posterior_predictive.nc
posterior_predictive.png
posterior_predictive_summary.csv
observed.csv
fitted.csv
fit_timeseries.png
fit_scatter.png
residuals.csv
residuals_timeseries.png
residuals_hist.png
residuals_vs_fitted.png
This stage is the in-sample or training-fit assessment. It uses the same data
the model was fit on and should not be read as the pipeline’s out-of-sample
validation layer.
35_holdout_validation
Main files:
validation_metadata.json
holdout_posterior_predictive.nc
holdout_predictive_summary.csv
holdout_predictive_report.json
holdout_observed.csv
holdout_fitted.csv
holdout_residuals.csv
holdout_timeseries.png
holdout_residuals_acf.png
The holdout summary and report include uncertainty-aware metrics such as
crps, bias, and fixed coverage columns for coverage_50, coverage_80,
and coverage_94.
This stage is optional. When validation is absent or disabled in YAML, the
directory still exists and the stage is marked skipped.
cre_decomposition_reconciliation.csv and
cre_decomposition_reconciliation.json for CRE
The CRE adjustment is retained on the baseline or non-incremental side of the
decomposition. Do not report it as an incremental media contribution.
50_diagnostics
Main files:
design_summary.csv
design_report.json
vif_report.csv
mcmc_summary.csv
mcmc_report.json
predictive_summary.csv
predictive_report.json
residual_diagnostics.csv
residuals_acf.png
diagnostics_report.csv
diagnostic_gates.resolved.yaml
diagnostics_summary.txt
chain_diagnostics.txt
The design-oriented files are raw input screening outputs. In particular,
diagnostics_report.csv labels the corresponding phase as
raw_input_screening rather than design.
diagnostic_gates.resolved.yaml records the profile name, source, inline
override status, exact boundary rule, and effective warn/fail values used for
each check. It is the audit trail for the run’s diagnostic decisions.
55_ai_diagnostics_advisor
Main files:
diagnostics_evidence.json
diagnostics_rules_summary.json
diagnostics_advisor_status.json
diagnostics_advisor_response.json, when an LLM advisor call succeeds
<provider>_diagnostics_response.raw.json, when an LLM advisor call succeeds
diagnostics_advisor_recommendations.md, when an advisor response is available
diagnostics_advisor_error.json, when the LLM call fails
This stage runs after 50_diagnostics by default for enabled ai_advisor
blocks. Set ai_advisor.diagnostics_review_enabled: false to skip it.
The diagnostics advisor evidence is restricted to anonymized channel aliases,
convergence counts, normalized predictive metrics, coverage metrics, and
scale-free design diagnostics. Raw target-scale fit errors are intentionally
excluded from the LLM payload.
The advisor separates computational reliability, raw-design identification
risk, predictive evidence, residual structure, prior robustness, and causal
identification. Its deterministic decision state is a floor: the LLM can make
the result stricter, but cannot soften a failed or warning gate.
60_response_curves
Main files:
saturation_curve.nc
saturation_curve_summary.csv
saturation_curve.png
forward_pass_contribution_curve.nc
forward_pass_contribution_curve_summary.csv
forward_pass_contribution_curve.png
adstock_curve.nc
adstock_curve_summary.csv
adstock_curve.png
These artefacts are intentionally different:
saturation_curve.* is the sampled saturation transformation on the scaled
channel axis, exported with original-scale contribution values for easier
reading. The PNG overlays that saturation-only curve against posterior mean
realised contributions.
forward_pass_contribution_curve.* is a full-model direct contribution
artefact. It rescales the observed historical spend path from 0% to 200%,
runs that spend through the fitted adstock and saturation path, and records
the resulting total channel contribution in original target units.
adstock_curve.* is the sampled carryover-weight profile for one impulse.
70_optimisation
This directory is present for every run, but the stage is skipped unless the
YAML config contains an optimization block.
Main files when the stage runs:
optimized_allocation.nc
optimized_allocation.csv
response_distribution.nc (optimised expected media contribution)
several PNG plots for allocation, contribution over time, response curves,
impact, bounds audit, and ROI or CPA
The two response datasets use identical posterior draws and outcome dates,
zero initial history, a shared time profile and no allocation or observation
noise. current means a reference allocation replay, not fitted historical
attribution. budget_optimisation.json records the comparison contract.
budget_impact.csv includes channel-level paired difference HDIs; total
difference HDIs come from summed paired draws, not sums of interval bounds.
Pipeline hdi_* fields now use ArviZ single-interval HDIs. Earlier pipeline
exports used equal-tailed bounds under these names. Recompute summaries from
retained draws; do not relabel historical files as corrected HDIs. Predictive
coverage remains a separate equal-tailed calculation.
80_interpretation
Main files:
evidence_inventory.md
interpretation_report.md
Both files currently contain the same evidence inventory: preceding stage
statuses, the holdout validation status and a reminder to review the retained
artefacts. They do not provide automated business conclusions. A completed
stage records successful report generation, not statistical qualification.
run_manifest.json
The manifest is the machine-readable index for the whole run.
Top-level fields include:
Field
Meaning
run_name
Effective run name
timestamp
UTC run timestamp
config_path
Original config path
output_dir
Run directory path
status
Overall run status
model_class
Identifies the model object prepared in Stage 00; does not certify graph completion
data
Basic dataset metadata
stages
Per-stage manifest records
warnings
Run-level warnings
error
Run-level failure payload when the pipeline aborts
data includes:
x_shape
y_length
target_column
x_columns
Stage records
Each stage record contains:
Field
Meaning
directory
Stage directory name
status
Current stage status
started_at
ISO timestamp when the stage started
finished_at
ISO timestamp when the stage finished
artifacts
Mapping of artefact labels to root-relative paths
warnings
Stage warnings
error
Error string when the stage fails
The artifacts mapping uses root-relative paths such as
20_model_fit/model.nc.
Stage statuses
Status
Meaning
pending
Stage has not started yet
running
Stage is currently running
completed
Stage finished successfully
skipped
Stage returned None intentionally
failed
Stage raised an exception
not_reached
A previous stage failed before this one ran
Common cases:
Stage 35 is skipped when validation is missing or disabled from YAML.
Stage 70 is skipped when optimization is missing from YAML.
Later stages become not_reached after the first failure.
Practical use
Use the run directory when you want:
a stable folder for downstream reporting
a machine-readable audit trail through run_manifest.json
stage-level links to artefacts without hard-coding filenames
When ai_advisor.enabled: true and the advisor proposes a valid patch, Stage
08 writes 08_ai_advisor/approval_request.yaml with status: pending. Review
the proposal, edit the request to status: approved, then run:
The retained runner is static, not plugin-based. To add a stage or integrate
custom status reporting, extend the existing runner surfaces instead of
bypassing them.
return a dict[str, str] of artefact labels to root-relative paths when the
stage succeeds
return None when the stage is intentionally skipped
raise an exception when the stage fails and should abort the run
The runner handles manifest updates around the stage call. Do not update
context.manifest directly from a normal stage implementation unless you are
changing core runner behaviour.
What is available in PipelineContext
PipelineContext gives each stage access to:
Field
Use it for
run_config
Runtime settings such as output root, seeds, and curve sample counts
raw_cfg
The loaded YAML config as a mutable mapping
X, y
Loaded dataset inputs
paths
Stage directories and manifest path
manifest
Current run manifest
model_kwargs
Effective sampler overrides passed into model build
Use context.paths.relative(path) when building the artefact mapping that the
stage returns. The manifest expects root-relative paths, not absolute paths.
fromabacus.pipeline.artifactsimportwrite_dataframedefrun_custom_stage(context):ifcontext.modelisNone:raiseValueError("Model has not been initialized before the custom stage.")stage_dir=context.paths.stage_dirs["custom"]output_path=stage_dir/"custom_summary.csv"frame=context.model.summary.total_contribution(output_format="pandas")write_dataframe(output_path,frame)return{"custom_summary":context.paths.relative(output_path),}
Optional stage pattern
If a stage should only run when a config block is present, follow the same
pattern as Stage 70:
Returning None is what marks the stage as skipped in the manifest.
Failure semantics
If your stage raises an exception:
the stage is marked failed
the run is marked failed
later pending stages are marked not_reached
run_pipeline(...) re-raises the exception
That means stage code should only catch exceptions when it can recover locally
and still produce a valid artefact set.
Adding structured reporting
If you want progress callbacks without changing the core stage code, implement a
PipelineReporter and pass it to run_pipeline(...).
The reporter protocol methods are:
on_pipeline_start(...)
on_stage_start(...)
on_stage_end(...)
on_pipeline_end(...)
on_pipeline_error(...)
This is the right extension point for:
notebooks or dashboards that want progress updates
lightweight orchestration wrappers
structured logging around pipeline runs
Consuming the manifest programmatically
The manifest is written after every stage transition, so external tools can
poll run_manifest.json during execution. Each update replaces the file
atomically from a completed temporary file in the same directory. Open the
manifest again for each poll to observe later snapshots; a previously opened
file can still refer to the older snapshot. Before initial publication the
manifest may be absent. If publication fails, the last complete snapshot
remains and the runner raises an error. A forcibly terminated process may
leave a temporary file, which is not a published manifest.
The runner owns manifest updates during execution. Atomic replacement does
not merge concurrent edits from other writers or make stage artefact writes
transactional.
Typical uses:
check whether the optimisation stage was skipped
discover stage artefact paths without hard-coding filenames
detect the first failed stage and its error message
Stage 35 is Abacus’s out-of-sample time-series validation layer.
It answers a narrower and more useful question than “does the model fit the
training data?”:
“If I refit the MMM on the earlier history only, can it still predict the last
blocked window reasonably well?”
For weekly MMM, that is usually a better stress test than a random split
because media carryover, seasonality, and trend all depend on time order.
Why Abacus uses a blocked tail holdout
Abacus reserves the final holdout_observations unique dates as the holdout
window, then fits a fresh model on the earlier dates only. The validation
pipeline does not reuse Stage 20 posterior state.
That means Stage 35 is checking whether the full model specification can
generalise forward in time, not whether the same fitted posterior can explain
the rows it already saw.
This is especially useful in MMM because:
adstock depends on lagged spend history
seasonality is time-ordered rather than exchangeable
marketing calendars often drift near the end of the sample
overfit specifications can look fine in-sample and fail on the final weeks
Reserve the last holdout_observations dates as the holdout window.
Fit a fresh model on the remaining earlier dates only.
Sample posterior predictive draws for the holdout rows.
Compute uncertainty-aware predictive metrics and residual diagnostics.
If include_last_observations: true, Abacus prepends the trailing lag history
needed for adstock carryover internally, then trims those prepended rows back
out of the returned holdout predictions. This matters whenever media effects
have memory.
Random seeds
Holdout posterior prediction explicitly receives the effective validation
fitting seed. Seed precedence is the same as for the refit: base YAML fit,
then runner sampler overrides, then validation.sampler. The metadata records
prediction_random_seed alongside the effective sampler_config. A zero seed
is valid; a missing or null effective seed leaves prediction unseeded and is
recorded as null.
Reusing a seed supports repeatability within the same environment and execution
settings; it does not promise identical draws across dependency versions or
backends. No independent-stream derivation is applied.
Why this stage exists when Stage 30 already exists
Stage 30 and Stage 35 answer different questions.
Stage 30 is an in-sample fit check. It uses the same rows the model was fit
on.
Stage 35 is an out-of-sample blocked holdout check. It uses future dates the
validation fit did not see.
If Stage 30 looks good and Stage 35 looks weak, that is a classic warning sign
of overfit, misspecification, or regime change.
Main artefacts
Stage 35 writes the following files under
results/<run_name>_<timestamp>_<random_suffix>/35_holdout_validation/:
The headline table is holdout_predictive_summary.csv.
The canonical predictive metric definitions
specify the formulas, observation aggregation and missing-value boundaries.
Point-error metrics
RMSE gives larger errors more weight than MAE. NRMSE and NMAE divide those
scores by the observed target range in the scored holdout, returning NaN
when that range is approximately zero. Compare scores on comparable windows;
range normalisation alone does not make different datasets comparable.
Use these as forecast-quality metrics, not causal-identification metrics.
Bias
bias is observed minus posterior predictive mean, averaged over observations:
positive bias means underprediction on average
negative bias means overprediction on average
For example, observed 12 and predicted mean 10 gives bias +2. Inspect persistent
bias alongside trend changes, omitted predictors and changes in the holdout
period. Its sign alone does not identify the cause.
CRPS
The continuous ranked probability score (CRPS) assesses the full predictive
distribution. Lower is better on the same evaluation set. Inspect it alongside
point errors, coverage and interval width; one score does not establish
calibration.
Coverage
Stage 35 reports coverage_50, coverage_80 and coverage_94, the observed
fractions inside equal-tailed posterior predictive intervals at those nominal
probabilities. They do not measure coverage of mmm.summary HDIs. See the
metric definitions
for the quantiles, endpoint inclusion and finite-observation denominator.
Compare coverage with its nominal probability alongside bias, interval width,
residual patterns, holdout size and dependence between observations. Low
coverage can reflect narrow intervals, prediction bias or distributional
change. High coverage does not by itself prove that intervals are too wide.
With eight finite aggregate observations, coverage changes in steps of 1/8.
Seven covered observations give 0.875, so exact agreement with 0.94 is
impossible. That result alone cannot establish or refute nominal 94%
calibration. Serial or panel dependence further limits the information in a
short holdout; panel rows do not automatically supply independent evidence.
How to read the plots
holdout_timeseries.png
This is the first plot to inspect.
Look for:
whether the observed series generally stays inside the predictive intervals
whether misses are isolated or systematically one-sided
whether the model misses turning points or holiday spikes
whether the predictive band width looks plausible relative to the volatility
of the target
Common interpretations:
repeated misses on the same side: likely bias
clustered misses: inspect bias, changing volatility and omitted time structure
wide bands: inspect predictive spread and its sensitivity to the specification;
width alone does not diagnose identification
holdout_residuals_acf.png
This checks whether the holdout residuals still contain serial structure.
Look for:
obvious positive autocorrelation across nearby lags
repeating seasonal patterns
long runs of same-sign residuals
If residual autocorrelation is strong, the model is usually still missing some
time structure such as:
seasonality
holiday dynamics
delayed media effects
structural breaks or time-varying baseline behavior
Practical rules of thumb
These are pragmatic MMM heuristics, not hard pass/fail thresholds.
Choosing the holdout window
For weekly MMM with roughly 1 to 2 years of data, 6 to 12 weeks is a
practical starting range.
8 weeks is a sensible default when you want enough tail signal without
throwing away too much history.
If the dataset is very short, a larger holdout can make the validation noisy
and can leave too little training history for stable estimation.
Abacus itself enforces that blocked holdout validation must still leave enough
training dates for the model to run; see
validation.py.
Comparing Stage 30 and Stage 35
Expect Stage 35 to be worse than Stage 30. That is normal.
Worry when the degradation is large or the direction changes materially.
If Stage 30 is excellent and Stage 35 is weak, suspect overfit or
misspecification before celebrating the in-sample fit.
Reading coverage
Report nominal probability, empirical coverage and the number of scored observations.
Investigate departures alongside bias and width; do not infer the cause from coverage alone.
Use additional comparable windows when feasible; one short holdout does not establish calibration.
Reading bias
A small nonzero bias is normal.
Large consistent bias over the holdout tail is a stronger warning sign than a
single noisy miss.
If bias keeps the same sign across several model variants, inspect trend,
holidays, and baseline structure before changing media priors.
Reading residual structure
White-noise-like residuals are what you want.
Visible residual runs or autocorrelation usually mean the model is still
missing systematic time variation.
Do not treat a decent RMSE as sufficient if residuals still show structure.
What Stage 35 does not tell you
Blocked holdout validation is valuable, but it is not a causal guarantee.
It does not prove:
that channel attribution is identified
that ROAS is unbiased
that the chosen priors are correct
that the model is safe for large budget reallocation on its own
It does tell you whether the specification can forecast a held-out tail window
coherently. That makes it an important diagnostic, but still only one part of
MMM model assessment.
Recommended workflow
For most weekly MMM work:
Run Stage 30 and Stage 35 together.
Compare in-sample and holdout metrics before changing the specification.
Use the same holdout window across candidate models so the comparison is
fair.
Prefer specifications that are stable across reasonable prior choices, not
just the one that scores best on a single holdout.
Treat Stage 35 as a forecasting sanity check alongside prior predictive
checks, posterior predictive checks, and substantive business review.
Common mistakes
Using a random split instead of a blocked time split
Reading Stage 30 as out-of-sample validation
Ignoring coverage and focusing only on RMSE
Using a holdout that is too long for the amount of history available
Repeatedly tuning the spec to one holdout window until the score looks good