Pipeline Runner

This section covers the structured abacus.pipeline runner: how it loads a config and dataset, executes the retained stage sequence, and writes reproducible run artefacts to disk.

Pages

  • Runner Overview - How run_pipeline(...) works, which stages run, and when the optimisation stage is skipped.
  • YAML Configuration - Which YAML keys the runner consumes and how they map to model build, data loading, holidays, and optimisation.
  • Blocked Holdout Validation - What Stage 35 does, how to configure it, and how to read the holdout metrics and plots.
  • CLI Reference - The thin python -m abacus.pipeline.runner interface and its supported flags.
  • Output Directory Schema - The run directory layout, manifest schema, stage statuses, and main artefacts.
  • Extending the Runner - How to add a stage or wire in reporting without bypassing the manifest and artifact helpers.

Subsections of Pipeline Runner

Runner Overview

Use the pipeline runner when you want a full disk-backed PanelMMM run instead of only an in-memory fit.

The runner loads a YAML config and a CSV dataset, builds the model, executes a fixed stage sequence, and writes each stage’s artefacts into a structured run directory. When validation is enabled, the runner performs a second train-window fit for the blocked holdout stage, so the run takes longer than a pure full-sample fit.

If you want a quick first run, start with Quickstart: Pipeline Runner.

Public entry points

The public Python API is:

  • abacus.pipeline.PipelineRunConfig
  • abacus.pipeline.run_pipeline
  • abacus.pipeline.PipelineRunResult

The thin CLI wraps the same code path:

python -m abacus.pipeline.runner --config path/to/config.yml

Basic Python example

from pathlib import Path

from abacus.pipeline import PipelineRunConfig, run_pipeline

result = run_pipeline(
    PipelineRunConfig(
        config_path=Path("data/demo/geo_panel/config.yml"),
        output_dir=Path("results"),
        run_name="geo_panel_baseline",
        prior_samples=10,
        draws=500,
        tune=500,
        chains=2,
        cores=2,
        random_seed=42,
        curve_samples=100,
        curve_points=100,
    )
)

print(result.run_dir)
print(result.manifest_path)

PipelineRunResult contains:

Field Meaning
run_dir The created run directory
manifest_path The path to run_manifest.json inside that directory

What the runner does

run_pipeline(...) performs these steps:

  1. Load the YAML config with load_yaml_config(...).
  2. Load X and y from CSV using load_pipeline_data(...).
  3. Merge CLI sampler overrides with YAML fit through build_model_kwargs(...).
  4. Create the output directory tree and initialise run_manifest.json.
  5. Run the retained stages in order, updating the manifest after every stage.

For named estimators, Stage 00 prepares the model and stores it in the shared PipelineContext; Stage 10 completes the deferred graph before prior predictive sampling. Configurations without a named estimator use the direct builder path, which can complete the graph in Stage 00. A populated model_class field identifies the model object, not graph completion.

Runner-only roots such as prior_sensitivity, ai_advisor, diagnostics, and validation stay on the pipeline context and are stripped before the public MMM builder validates the model YAML.

Stage order

This is the canonical stage sequence. The runner uses a fixed stage list; optional stages retain their place and record a skipped status when disabled. Stage 80 inventories evidence and stage statuses. It does not generate automated business conclusions or certify decision readiness.

Stage key Directory Purpose Optional
metadata 00_run_metadata Prepare the model and write resolved config and dataset metadata No
prior_sensitivity 05_prior_sensitivity Write resolved prior-sensitivity scenario configs and manifests Yes
ai_advisor 08_ai_advisor Write privacy-safe deterministic and optional LLM guidance artifacts Yes
preflight 10_pre_diagnostics Complete any deferred graph; draw and plot prior predictions No
fit 20_model_fit Fit the model, save InferenceData, write trace and summary No
assessment 30_model_assessment In-sample posterior predictive checks, fitted values, residual outputs No
validation 35_holdout_validation Blocked holdout scoring on a train-window refit Yes
decomposition 40_decomposition Contribution tables and decomposition plots No
diagnostics 50_diagnostics Raw input screening, MCMC, predictive, and residual diagnostics No
ai_diagnostics_advisor 55_ai_diagnostics_advisor Write privacy-safe LLM diagnostics review artifacts for enabled AI advisor runs Yes
curves 60_response_curves Saturation-only, forward-pass direct contribution, and adstock curve artefacts No
optimisation 70_optimisation Budget optimisation artefacts Yes
interpretation 80_interpretation Inventory retained evidence for analyst review No

The prior-sensitivity stage is marked skipped when the YAML config does not contain prior_sensitivity or it is disabled. The AI advisor stage follows the same convention for ai_advisor. The diagnostics advisor stage runs by default for enabled ai_advisor blocks and is marked skipped only when ai_advisor is absent, disabled, or has diagnostics_review_enabled: false. The validation stage is marked skipped when the YAML config does not contain validation or it is disabled. The optimisation stage is also optional; it returns None and is marked skipped when the YAML config does not contain an optimization block.

See Output Directory Schema for the stage folders and artefact layout.

Data and model assumptions

The retained runner is designed around PanelMMM.

  • The flow-oriented public YAML is expected to describe a PanelMMM.
  • The data loader reads CSV only.
  • Later stages call PanelMMM plotting, summary, diagnostics, and optimisation methods directly.

If you need the exact YAML keys, see YAML Configuration.

PipelineRunConfig

PipelineRunConfig controls runtime settings that sit outside the YAML model specification.

Field Purpose
config_path YAML file to load
output_dir Root directory under which the run directory is created
run_name Optional run-name override; otherwise the config filename stem
dataset_path Optional combined dataset CSV override
x_path, y_path Optional feature and target CSV overrides
holidays_path Optional holiday CSV override
target_column Target column name used during CSV loading
prior_samples Number of prior predictive samples for Stage 10
draws, tune, chains, cores, random_seed Sampler overrides merged onto YAML fit
curve_samples, curve_points Curve sampling settings for Stage 60

The draws, tune, chains and cores overrides do not necessarily bound Stage 35: explicit validation.sampler values take precedence for the refit. Use the bounded software smoke for an execution check with validation explicitly skipped.

Only sampler settings are merged into model construction. Other overrides are used by the runner itself during data loading, holiday resolution, diagnostics reporting, and output setup.

Run directory naming

The runner creates the run directory as:

<output_dir>/<effective_run_name>_<YYYYMMDD_HHMMSS>_<random_suffix>

The timestamp is generated in UTC. Each invocation exclusively allocates a new run directory, including simultaneous invocations with the same name and timestamp. The random suffix is opaque; use PipelineRunResult.run_dir and manifest_path instead of reconstructing paths from the name and timestamp. Existing runs are never reused or resumed by this allocator.

A run name must be a non-empty filename component, without path separators. Set output_dir to choose its parent. Allocated run directories have owner-only permissions on POSIX (0700); manifest files have mode 0600.

All stage directories are created up front, even if a later stage is skipped or the run aborts. An allocation or stage-directory setup error propagates; a partially initialised new run directory may remain for inspection.

Manifest updates are published by replacing the previous file with a completed temporary file in the same directory. Readers see a complete old or new snapshot. A publication failure propagates and leaves the previous snapshot intact, which may therefore be stale. This does not make stage artefacts transactional or provide power-loss durability.

Failure and skip behaviour

If a stage raises an exception:

  • the current stage is marked failed
  • the run manifest is marked failed
  • all still-pending later stages are marked not_reached
  • run_pipeline(...) re-raises the exception

If a stage returns None:

  • the stage is marked skipped
  • the manifest warning records that no configuration was supplied for that optional stage

Reporter hook

run_pipeline(...) accepts an optional reporter that implements the PipelineReporter protocol.

The reporter can observe:

  • pipeline start
  • stage start
  • stage end
  • pipeline end
  • pipeline failure

See Extending the Runner for the callback contract.

YAML Configuration

The pipeline runner reads the same YAML model specification used by build_mmm_from_yaml(...), then adds a small set of runner-specific conventions for data loading, prior-sensitivity planning, optional AI advisor guidance, optional blocked holdout validation, and Stage 70 optimisation.

This page documents the keys that the runner actually consumes.

Root keys

Key Required Used for
data Usually Resolve dataset paths when you do not pass dataset_path, x_path, or y_path through PipelineRunConfig
target Yes Define the target column and business target type
estimator No Declare a named estimator preset; time_series, fe, and cre are currently released
dimensions No Declare panel-dimension columns such as geo or brand
media Yes Define channel/control columns and transform types
scaling No Configure target/channel scaling rules
effects No Append additive effects in YAML order before build_model(...)
priors No Override model-level priors and prefixed transform priors
fit No Default sampler settings for Stage 20 fitting
holidays No Add holiday events before model build
original_scale_vars No Add original-scale contribution variables before fitting
inference_data No Attach existing InferenceData when the file exists
prior_sensitivity No Write a pre-fit scenario plan for prior robustness checks
ai_advisor No Write privacy-safe AI advisor guidance before model fitting
validation No Enable optional Stage 35 blocked holdout validation
optimization No Enable Stage 70 budget optimisation
diagnostics No Override Stage 50 runner diagnostics thresholds

Minimal runner config

data:
  dataset_path: dataset.csv
  date_column: date

target:
  column: revenue
  type: revenue

estimator:
  type: time_series

media:
  channels: [channel_1, channel_2]
  adstock:
    type: geometric
    l_max: 4
  saturation:
    type: logistic

fit:
  draws: 1000
  tune: 1000
  chains: 4
  cores: 4
  random_seed: 42

Relative paths in YAML are resolved relative to the YAML file’s directory.

diagnostics is runner-only. The structured pipeline reads it, but build_mmm_from_yaml(...) still validates only the public MMM model schema.

prior_sensitivity and ai_advisor are also runner-only. They are consumed by the structured pipeline before model fitting and stripped before the public MMM YAML builder validates the model specification.

validation is also runner-only. The structured pipeline reads it for Stage 35 blocked holdout scoring, but the public MMM YAML builder never sees it.

Core modeling blocks

The runner always builds a PanelMMM, so the public YAML no longer exposes a model.class field. Instead, it reads:

  • data.date_column
  • target.column
  • target.type
  • media.channels
  • media.controls, if any
  • estimator, for a named preset
  • dimensions.panel, if any
  • media.adstock
  • media.saturation
  • fit

estimator

The released named single-series contract is:

estimator:
  type: time_series

It requires one observation per date and no panel unit. It builds the same single-series graph as the established configuration with no dimensions.panel for the same configuration. With the default model settings, this means one global intercept and shared media, control, adstock, saturation, and residual parameters. Other explicit single-series model options retain their established behaviour; the estimator declaration does not silently override them.

The released fixed-effects contract is:

estimator:
  type: fe
  unit: geo
  estimability:
    within_variation_share_warning: 0.05
    max_vif_warning: 20
    condition_number_warning: 30

It accepts one unit column and uses an exact within-unit orthonormal-contrast likelihood. Unit intercepts are absorbed. Media and control slopes, adstock, saturation, and residual scale are shared across units. The FE preset does not support common time effects, annual seasonality, custom additive effects, or time-varying parameters. See Fixed-effects Estimator for the estimability checks and interpretation limits.

The released correlated-random-effects contract is:

estimator:
  type: cre
  unit: geo
  estimability:
    within_variation_share_warning: 0.05
    max_vif_warning: 20
    condition_number_warning: 30
    minimum_between_residual_df: 2
    posterior_diagnostic_draws: 50

It accepts one balanced unit panel. Media and control slopes, geometric adstock, logistic saturation, and residual scale are shared across units. The graph uses an exact marginal Gaussian random-intercept likelihood and adds centred unit means of the transformed media basis and eligible time-varying controls. Common time effects, seasonality, custom effects, calibration, optimisation and fixed-budget scenario optimisation are not supported. Prediction and historical/manual scenarios require all fitted units and reject unseen units and unit subsets. Manual CRE scenarios retain the fitted training-period Mundlak summaries rather than recomputing them from planned spend. See Correlated-random-effects Estimator for the estimability and interpretation limits.

The re declaration validates as typed configuration but remains release-gated. It fails before graph construction and does not fall back to the advanced panel-dimension surface.

Do not combine estimator with dimensions.panel. Abacus rejects the mixed declaration rather than guessing which semantics you intended.

data

The runner loads data before building the model. It supports two CSV layouts.

Combined dataset

data:
  dataset_path: "dataset.csv"

The runner reads the CSV, removes the target column from X, and uses that column as y.

Separate feature and target files

data:
  x_path: "X.csv"
  y_path: "y.csv"

When loading y_path:

  • if the configured target column exists, the runner uses that column
  • otherwise, if the file has exactly one column, the runner uses that column and renames it to the target name

Target column resolution

The runner resolves the target column in this order:

  1. PipelineRunConfig.target_column or CLI --target-column
  2. target.column
  3. "y"

Use the CLI override only when you want to change how the runner reads the CSV. Keep it consistent with target.column in YAML.

fit

fit controls Stage 20 fitting because the fit stage calls:

context.model.fit(X=context.X, y=context.y, progressbar=False)

The runner merges these CLI or PipelineRunConfig overrides onto the YAML fit block when they are provided:

  • draws
  • tune
  • chains
  • cores
  • random_seed

The public YAML schema currently supports these fit keys:

  • draws
  • tune
  • chains
  • cores
  • random_seed
  • target_accept
  • progressbar
  • compute_convergence_checks

Unknown fit keys are rejected when the YAML is loaded.

effects

effects is an optional list of additive effect specifications:

effects:
  - type: linear_trend
    prefix: trend
    n_changepoints: 8
  - type: weekly_fourier
    order: 3

The builder appends each effect to model.mu_effects in YAML order before calling build_model(...).

holidays

The holidays block is optional.

Supported keys used by the builder include:

Key Meaning
path Holiday CSV path
enabled Set to false to disable holiday loading
prefix Prefix for generated holiday effect coordinates
mode Holiday handling mode: event, pooled_control, or prophet_component
countries Country filter for catalogue-style holiday CSV input

Example:

holidays:
  mode: prophet_component
  path: "../../data/holidays.csv"
  prefix: "holiday"
  countries: "UK"

The CLI or PipelineRunConfig.holidays_path overrides holidays.path.

If you omit both path and the override but still configure holidays, Abacus falls back to the bundled abacus.data:holidays.csv.

Country-selection rules:

  • time-series configs default to US when holidays.countries is omitted
  • geo-panel configs must declare holidays.countries explicitly
  • geo-panel configs must provide multiple countries, for example ["UK", "FR", "DE"]

If you provide a catalogue-style holiday CSV, Abacus only creates holiday effects for the countries listed in holidays.countries.

Holiday modes:

  • event creates one latent holiday/event effect per holiday row, which is why posterior summaries include terms like holiday_effect_size[...].
  • pooled_control creates one pooled binary holiday regressor over time and estimates a single shared holiday coefficient. This is useful when you want a strict calendar-only single holiday term instead of one parameter per holiday.
  • prophet_component fits Prophet on the training target with the configured holiday calendar, extracts the continuous holidays component, and uses that single smoothed series inside the MMM as one holiday term. For panel models, Abacus fits one Prophet holiday component per panel series and filters the holiday calendar by geo when that dimension is present.

For holiday effects, each model date labels the start of its observed period. Abacus assigns an inclusive holiday date range to every model period it overlaps. For example, on W-MON data, a Wednesday or Sunday holiday is assigned to the Monday date that starts that week. Daily data retains its existing date-by-date assignment.

Default behavior:

  • configs default to prophet_component
  • use event explicitly when you want one latent holiday effect per holiday row

Current limitation:

  • pooled_control currently supports only single-country, non-geo models.
  • prophet_component requires exactly one holiday country unless the model has a geo dimension, in which case it can route multiple holiday countries to the matching geo-level panel series.

original_scale_vars

Use original_scale_vars when you want specific contribution variables to be available on the original target scale:

original_scale_vars:
  - channel_contribution
  - y

The builder applies these through model.add_original_scale_contribution_variable(...) before fitting.

inference_data

inference_data.path is passed through to the YAML builder. If the file exists, Abacus attaches that InferenceData when the build completes: in Stage 10 for named estimators with a deferred graph, or Stage 00 for the direct builder path. See the runner lifecycle.

Important: the structured runner still executes Stage 20 and fits the model again. inference_data.path does not currently skip fitting.

prior_sensitivity

Use the optional prior_sensitivity block when you want the runner to write a pre-fit prior scenario plan. This stage does not fit every scenario. It creates resolved scenario configs that can be reviewed, approved, and run deliberately.

Conservative generated plan:

prior_sensitivity:
  enabled: true
  scenario_policy: conservative_mmm
  reference: reference

Manual plan:

prior_sensitivity:
  enabled: true
  scenario_policy: manual
  reference: reference
  scenarios:
    reference:
      description: Current approved prior specification.
    tighter_media_effect:
      description: Lower media-effect amplitude on the scaled target space.
      overrides:
        media.saturation.priors.beta:
          distribution: HalfNormal
          sigma: 0.5
          dims: ["channel"]

Supported keys:

Key Meaning
enabled Set to true to write Stage 05 prior-sensitivity artifacts
scenario_policy manual for declared scenarios or conservative_mmm for generated relative scenarios
reference Scenario name for the unchanged reference config
scenarios Optional manual scenario declarations
allow_model_structure_overrides Required before scenarios can change transform structure such as media.adstock.l_max

Scenario names are slugs such as reference, longer_memory, or tighter_media_effect. Avoid names such as baseline; in MMM, baseline has a model meaning and should not be overloaded as a scenario label.

Allowed override paths are intentionally narrow:

  • media.adstock.priors.*
  • media.saturation.priors.*
  • priors.*
  • selected transform-structure paths such as media.adstock.l_max, only when allow_model_structure_overrides: true

The stage writes both a human-readable manifest and an LLM-safe manifest. Use the LLM-safe file when passing scenario context to an external model because it aliases override paths and avoids free-text descriptions.

ai_advisor

Use the optional ai_advisor block when you want privacy-safe, evidence-grounded modelling guidance from deterministic rules and, optionally, OpenAI or OpenRouter. The advisor proposes controlled tests. It does not approve a model, establish causal identification, or apply a config patch to the run config.

ai_advisor:
  enabled: true
  provider: openrouter
  mode: autopilot
  privacy: anonymized_relative
  approval: file_based
  write_outputs: true
  llm_enabled: true
  diagnostics_review_enabled: true
  openai_model: gpt-5-mini
  openai_timeout_seconds: 60
  openrouter_model: openai/gpt-5.2
  openrouter_timeout_seconds: 60

Supported keys:

Key Meaning
enabled Set to true to write Stage 08 advisor artifacts
provider openai or openrouter
mode autopilot; the advisor prioritizes concise recommendations and approval-ready options
privacy anonymized_relative; raw channel names and raw business values are excluded from the LLM payload
approval file_based; proposed config changes are written as files for user approval
write_outputs Set to false to disable artifact writes even when the block is enabled
llm_enabled Set to false to run deterministic privacy/rule checks without an LLM call
diagnostics_review_enabled Defaults to true; set to false to skip the post-fit 55_ai_diagnostics_advisor LLM review after structured diagnostics
openai_model OpenAI model name used for the advisor call
openai_timeout_seconds Request timeout for the OpenAI call
openrouter_model OpenRouter model name used for the advisor call
openrouter_timeout_seconds Request timeout for the OpenRouter call

The pipeline reads OPENAI_API_KEY or OPENROUTER_API_KEY from the process environment based on provider. For local development, an untracked repo-root .env file is also supported. Do not commit API keys.

The advisor stages complete even if an LLM call fails. In that case they write an error artifact and the rest of the pipeline can continue. Deterministic rules provide a minimum decision state: an LLM may make the state stricter, but cannot override a failed gate or weak-identification warning with a more favourable conclusion.

When the advisor proposes a valid config patch, Stage 08 writes:

  • config_patch_proposal.yaml
  • approval_request.yaml

The proposal format is deliberately narrow:

overrides:
  media.saturation.priors.beta:
    distribution: HalfNormal
    sigma: 0.5
    dims: ["channel"]

To approve it, edit approval_request.yaml so status: approved, then run:

python -m abacus.pipeline.approval \
  --approval-request results/<run>/08_ai_advisor/approval_request.yaml \
  --approved-by "model owner"

The approval command writes approved_config.resolved.yaml and approval_record.yaml beside the advisor artifacts. It does not mutate the source YAML config.

By default, an enabled ai_advisor block also runs 55_ai_diagnostics_advisor after structured diagnostics. That post-fit advisor uses anonymized channel aliases, convergence counts, normalized predictive metrics, coverage metrics, and scale-free design diagnostics. It intentionally excludes raw target-scale fit errors from the LLM payload. Set diagnostics_review_enabled: false to run only the pre-fit advisor.

optimization

Add an optimization block when you want Stage 70 to run. If this block is absent, Stage 70 is marked skipped.

The YAML builder validates this block when the config is loaded. start_date and end_date are always required, and you must provide exactly one of:

  • optimization.budget for the preferred user-facing budget spec
  • optimization.total_budget for the legacy per-period budget input

Unknown top-level optimization keys are rejected.

Preferred example:

optimization:
  start_date: "2024-11-11"
  end_date: "2025-01-27"
  budget:
    mode: relative
    value: 1.10
    basis: reference_window_total

Optional keys read by Stage 70:

Key Default Meaning
budget None Preferred user-facing budget spec: absolute or relative
total_budget None Legacy per-period budget input kept for backward compatibility
response_variable total_media_contribution_original_scale Optimisation objective variable
budget_distribution_over_period None Time weights over the optimisation window
budget_bounds Derived or default Explicit spend bounds
minimize_kwargs None Options forwarded to the existing SciPy minimiser contract; defaults remain SLSQP, ftol=1e-9, maxiter=1000
spend_constraint_lower 0.3 when deriving bounds Relative lower bound around scaled reference spend
spend_constraint_upper 0.3 when deriving bounds Relative upper bound around scaled reference spend
default_constraints true Whether to add the default equality budget constraint
noise_level 0.0 Must be zero for the deterministic allocation comparison
include_last_observations false Must be false, matching the optimiser initial history
include_carryover true Must be true, matching the optimiser response horizon

Stage 70 rejects other values for these three comparison options. Both plans use the same posterior draws, time profile, zero initial history and full carryover horizon. The current row replays reference-window spend totals under that profile; it is not historical attribution. The comparison summarises expected media contribution, without observation noise.

Budget spec modes:

  • budget.mode: absolute budget.value is total spend over the full optimisation horizon.
  • budget.mode: relative budget.value is a multiplier on the chosen basis.
  • budget.basis: reference_window_total Abacus resolves the budget against the same reference-window total spend it already uses for current-plan comparison and default bound derivation.

Important budget-unit note

The preferred optimization.budget block uses total horizon spend. Stage 70 converts that to the wrapper’s per-period contract internally before calling PanelBudgetOptimizerWrapper.optimize_budget(...).

The legacy optimization.total_budget field is still supported, but it keeps the old wrapper-facing per-period spend contract.

See Budget Optimisation.

Xarray-like optimisation values in YAML

For panel bounds or time distributions, use the xarray-like mapping shape that Stage 70 expects:

optimization:
  start_date: "2025-02-03"
  end_date: "2025-02-24"
  budget:
    mode: absolute
    value: 100000.0
  budget_distribution_over_period:
    values:
      - [[0.25, 0.25], [0.25, 0.25]]
      - [[0.25, 0.25], [0.25, 0.25]]
      - [[0.25, 0.25], [0.25, 0.25]]
      - [[0.25, 0.25], [0.25, 0.25]]
    dims: ["date", "geo", "channel"]
    coords:
      date: [0, 1, 2, 3]
      geo: ["UK", "FR"]
      channel: ["channel_1", "channel_2"]

Time profiles must include every budget dimension with unique labels matching the model exactly. Labels and dimensions can be reordered; the optimiser aligns them before use. Fractions must be finite and non-negative and sum to one along date for every budget cell. Do not omit labels or use negative weights to balance a column total.

The same shape works for budget_bounds, but with an additional "bound" dimension containing "lower" and "upper".

diagnostics

Stage 50 resolves a complete, versioned decision-gate profile and writes it to 50_diagnostics/diagnostic_gates.resolved.yaml. The packaged default is abacus/pipeline/diagnostic_gates.default.yaml. A documented copy is available at examples/diagnostic_gates.team.yaml. Copy that file when your team needs a governed profile with different thresholds; keep the source profile in version control with the model configuration.

Use gates_file to select that profile. Relative paths are resolved from the model YAML file. Optional inline thresholds take precedence over the selected profile and are recorded in the resolved artifact.

diagnostics:
  gates_file: diagnostic_gates.team_v1.yaml
  thresholds:
    design_max_vif:
      warn: 10.0
      fail: 20.0
    mcmc_max_rhat:
      warn: 1.02
      fail: 1.08

Supported threshold keys:

  • design_max_vif
  • design_condition_number
  • mcmc_divergence_count
  • mcmc_max_rhat
  • mcmc_min_ess_bulk
  • mcmc_bfmi_min
  • bayesian_pareto_k_max
  • predictive_nrmse
  • residual_ljung_box_p
  • residual_max_abs_acf

Validation rules:

  • upper-bound checks require warn <= fail
  • lower-bound checks require warn >= fail
  • equality triggers the relevant warn or fail boundary; the zero-divergence gate is the explicit exception, where zero passes and any positive count fails
  • a selected gate file must be schema version 1 and define every supported gate
  • omit the block entirely to use the packaged default profile

These gates classify available diagnostic evidence. Passing them does not prove parameter identification, prior robustness, model validity, or causal identification. In particular, VIF and condition number are raw-design screens. A warning indicates weak-identification risk and should trigger controlled reparameterisation or prior-sensitivity runs. A clean screen only means that no material warning was detected by those checks.

This block affects only the structured runner. It is stripped before Stage 00 model preparation so the public MMM YAML schema remains unchanged.

validation

Use the optional validation block when you want Stage 35 blocked holdout scoring. This is the runner’s out-of-sample tail check: Abacus refits a clean model on the earlier dates and scores only the final blocked window.

validation:
  enabled: true
  holdout_observations: 8
  include_last_observations: true
  coverage_levels: [0.5, 0.8, 0.94]
  sampler:
    draws: 500
    tune: 500
    chains: 2
    cores: 2
    random_seed: 42

Supported keys:

Key Meaning
enabled Set to false to skip Stage 35 while keeping the stage in the manifest
holdout_observations Number of unique dates to reserve for the blocked holdout window
include_last_observations Keep lag history for carryover-sensitive holdout scoring
coverage_levels Coverage levels reported in Phase 10; use the fixed 50, 80, and 94 percent defaults
sampler Optional validation-only sampler overrides for the train-window refit

Validation settings are merged in this order: YAML fit, runner sampler overrides, then validation.sampler. The last value wins. Reducing CLI draws or tuning does not override an explicit validation budget.

Phase 10 reports coverage as coverage_50, coverage_80, and coverage_94. Keep those defaults unless the implementation and tests are updated together.

The validation stage builds a clean train-window model for holdout scoring and ignores inference_data.path so the refit does not inherit attached posterior state from the main model build.

For a full explanation of why the split is blocked, how to read crps and coverage, and rules of thumb for weekly MMM, see Blocked Holdout Validation.

Override precedence

For the runner, precedence is:

Setting Higher precedence Lower precedence
Combined dataset path dataset_path / --dataset-path data.dataset_path
Split CSV paths x_path, y_path / --x-path, --y-path data.x_path, data.y_path
Holiday CSV path holidays_path / --holidays-path holidays.path
Sampler settings PipelineRunConfig or CLI overrides fit
Target column for CSV loading target_column / --target-column target.column, then "y"
Diagnostics thresholds diagnostics.thresholds retained Stage 50 defaults

Common pitfalls

  • Using Parquet paths in the pipeline data block. The runner data loader reads CSV only.
  • Providing only one of data.x_path or data.y_path.
  • Mixing the preferred horizon-based optimization.budget block with the legacy per-period optimization.total_budget field.
  • Assuming diagnostics is part of the public MMM builder schema. It is a runner-only block.
  • Assuming inference_data.path skips Stage 20 fitting. It does not.
  • Forgetting that relative paths are resolved from the YAML file directory, not from the shell working directory.

For large original-currency objectives, record an explicitly justified numerical tolerance in optimization.minimize_kwargs, for example {options: {ftol: 0.000001, maxiter: 1000}}. Verify feasibility and objective accuracy independently; successful termination alone is not statistical validation. Abacus does not automatically retry failed solvers with looser tolerances. The underlying Python wrapper accepts the same minimize_kwargs.

Output Directory Schema

Each pipeline run creates a timestamped directory under the configured output_dir:

<output_dir>/<run_name>_<YYYYMMDD_HHMMSS>_<random_suffix>

The timestamp is generated in UTC. A random suffix and exclusive directory creation keep same-name, same-time runs separate. Use the returned run path; do not infer a path from the timestamp. Existing saved runs remain readable. On POSIX, newly allocated run directories have mode 0700 and manifests have mode 0600.

The runner creates every stage directory up front, then publishes complete run_manifest.json snapshots by atomic replacement as stages start, complete, skip, or fail. Each run has its own manifest and artefacts.

Directory tree

results/
  geo_panel_baseline_20260308_153000_a8k2m7q1/
    run_manifest.json
    00_run_metadata/
    05_prior_sensitivity/
    08_ai_advisor/
    10_pre_diagnostics/
    20_model_fit/
    30_model_assessment/
    35_holdout_validation/
    40_decomposition/
    50_diagnostics/
    55_ai_diagnostics_advisor/
    60_response_curves/
    70_optimisation/
    80_interpretation/
    scenario_planner/
      recipes/
        <recipe>_<timestamp>/

scenario_planner/recipes/ is a post-fit evidence area, not a pipeline stage. It appears only after a user evaluates a retained scenario recipe. Recipe evaluation does not mutate run_manifest.json or refit the model.

Stage directories

Stage Directory Typical artefacts
metadata 00_run_metadata resolved config, model metadata, and estimator contract
prior_sensitivity 05_prior_sensitivity scenario configs, human manifest, and LLM-safe manifest
ai_advisor 08_ai_advisor privacy-safe evidence, rule summary, optional LLM response, and optional patch proposal
preflight 10_pre_diagnostics prior predictive and estimator-specific design evidence
fit 20_model_fit fitted model, trace, posterior summary, and CRE post-fit screen
assessment 30_model_assessment in-sample posterior predictive checks and residual outputs
validation 35_holdout_validation blocked holdout scoring, uncertainty-aware metrics, and residual diagnostics
decomposition 40_decomposition contribution CSVs, CRE reconciliation, and decomposition plots
diagnostics 50_diagnostics raw input screening, MCMC, predictive, and residual diagnostic reports
ai_diagnostics_advisor 55_ai_diagnostics_advisor privacy-safe post-fit diagnostics evidence and optional LLM guidance
curves 60_response_curves saturation-only, forward-pass direct contribution, and adstock NetCDF, summaries, and plots
optimisation 70_optimisation allocation, response, optimisation summary, and bounds audit artefacts
interpretation 80_interpretation evidence inventory and analyst review reminder

See Runner Overview for the stage order and optionality.

Post-fit scenario recipe bundles

Each recipe evaluation creates a new directory under scenario_planner/recipes/. It contains the resolved request, validation evidence, fitted estimator manifest, allocation tables, posterior media-contribution summaries with 94% highest-density intervals, a versioned dashboard payload, and a SHA-256 artefact manifest.

See Comparison Outputs for the complete file contract and interpretation boundary.

Main artefacts by stage

00_run_metadata

Main files:

  • a copy of the original config under its source filename
  • config.original.yaml
  • config.resolved.yaml
  • session_info.txt
  • dataset_metadata.json
  • model_metadata.json
  • data_dictionary.csv
  • design_matrix_manifest.csv
  • spec_summary.csv
  • estimator_summary.txt and estimator_manifest.yaml for named estimators
  • estimator_estimability.csv for the resolved estimator screen summary
  • holiday_feature_manifest.csv when holidays are configured

config.resolved.yaml normalises configured data and holiday paths to absolute paths and records the effective sampler configuration on the model.

05_prior_sensitivity

Main files:

  • scenario_manifest.yaml
  • llm_safe_scenario_manifest.yaml
  • <scenario_name>/config.resolved.yaml for each generated or declared scenario

This stage is optional. When prior_sensitivity is absent or disabled in YAML, the directory still exists and the stage is marked skipped.

scenario_manifest.yaml is the local, human-readable manifest. It can include scenario descriptions and raw override paths. Use llm_safe_scenario_manifest.yaml when passing scenario context to an external LLM because it aliases override paths and avoids free-text scenario prose.

08_ai_advisor

Main files:

  • evidence.json
  • rules_summary.json
  • advisor_status.json
  • advisor_response.json, when an LLM advisor call succeeds
  • <provider>_response.raw.json, when an LLM advisor call succeeds
  • advisor_recommendations.md, when an advisor response is available
  • config_patch_proposal.yaml, when the advisor proposes a YAML patch
  • approval_request.yaml, when the proposed patch is valid and ready for file-based approval
  • advisor_error.json, when the LLM call fails
  • config_patch_error.json, when an advisor patch is invalid or unsafe

This stage is optional. When ai_advisor is absent or disabled in YAML, the directory still exists and the stage is marked skipped.

The advisor evidence is privacy-safe by design: it aliases media channels, prior paths, and prior specs, and excludes raw business values. Deterministic rule checks always run before the optional LLM call. If those rules mark the payload as unsafe, the LLM call is blocked and the status artifact records the reason.

Patch proposals are validated before an approval request is written. Supported patches use an overrides: mapping with safe prior paths only. After a user sets approval_request.yaml to status: approved, the file-based approval CLI writes approved_config.resolved.yaml and approval_record.yaml in this directory.

10_pre_diagnostics

Main files:

  • prior_predictive.nc
  • prior_predictive.png
  • fixed_effects_estimability.csv and fixed_effects_estimability.json for FE
  • cre_structural_estimability.json, cre_reference_estimability.json, and cre_reference_estimability_features.csv for CRE

20_model_fit

Main files:

  • model.nc
  • trace.png
  • posterior_summary.csv
  • cre_postfit_estimability.json for CRE

posterior_summary.csv is intentionally compact. It summarizes structural posterior parameters such as adstock, saturation, seasonality, holiday, and likelihood terms, and omits per-date deterministic series like weekly channel contributions or fitted paths. Use the assessment and decomposition stages for time-indexed fitted or contribution outputs.

30_model_assessment

Main files:

  • posterior_predictive.nc
  • posterior_predictive.png
  • posterior_predictive_summary.csv
  • observed.csv
  • fitted.csv
  • fit_timeseries.png
  • fit_scatter.png
  • residuals.csv
  • residuals_timeseries.png
  • residuals_hist.png
  • residuals_vs_fitted.png

This stage is the in-sample or training-fit assessment. It uses the same data the model was fit on and should not be read as the pipeline’s out-of-sample validation layer.

35_holdout_validation

Main files:

  • validation_metadata.json
  • holdout_posterior_predictive.nc
  • holdout_predictive_summary.csv
  • holdout_predictive_report.json
  • holdout_observed.csv
  • holdout_fitted.csv
  • holdout_residuals.csv
  • holdout_timeseries.png
  • holdout_residuals_acf.png

The holdout summary and report include uncertainty-aware metrics such as crps, bias, and fixed coverage columns for coverage_50, coverage_80, and coverage_94.

This stage is optional. When validation is absent or disabled in YAML, the directory still exists and the stage is marked skipped.

For interpretation guidance and practical rules of thumb, see Blocked Holdout Validation.

40_decomposition

Main files:

  • waterfall_components_decomposition.png
  • weekly_media_contribution.png
  • channel_contributions.csv
  • baseline_contributions.csv
  • mean_contributions_over_time.csv
  • cre_adjustment_contributions.csv for CRE
  • cre_decomposition_reconciliation.csv and cre_decomposition_reconciliation.json for CRE

The CRE adjustment is retained on the baseline or non-incremental side of the decomposition. Do not report it as an incremental media contribution.

50_diagnostics

Main files:

  • design_summary.csv
  • design_report.json
  • vif_report.csv
  • mcmc_summary.csv
  • mcmc_report.json
  • predictive_summary.csv
  • predictive_report.json
  • residual_diagnostics.csv
  • residuals_acf.png
  • diagnostics_report.csv
  • diagnostic_gates.resolved.yaml
  • diagnostics_summary.txt
  • chain_diagnostics.txt

The design-oriented files are raw input screening outputs. In particular, diagnostics_report.csv labels the corresponding phase as raw_input_screening rather than design.

diagnostic_gates.resolved.yaml records the profile name, source, inline override status, exact boundary rule, and effective warn/fail values used for each check. It is the audit trail for the run’s diagnostic decisions.

55_ai_diagnostics_advisor

Main files:

  • diagnostics_evidence.json
  • diagnostics_rules_summary.json
  • diagnostics_advisor_status.json
  • diagnostics_advisor_response.json, when an LLM advisor call succeeds
  • <provider>_diagnostics_response.raw.json, when an LLM advisor call succeeds
  • diagnostics_advisor_recommendations.md, when an advisor response is available
  • diagnostics_advisor_error.json, when the LLM call fails

This stage runs after 50_diagnostics by default for enabled ai_advisor blocks. Set ai_advisor.diagnostics_review_enabled: false to skip it.

The diagnostics advisor evidence is restricted to anonymized channel aliases, convergence counts, normalized predictive metrics, coverage metrics, and scale-free design diagnostics. Raw target-scale fit errors are intentionally excluded from the LLM payload.

The advisor separates computational reliability, raw-design identification risk, predictive evidence, residual structure, prior robustness, and causal identification. Its deterministic decision state is a floor: the LLM can make the result stricter, but cannot soften a failed or warning gate.

60_response_curves

Main files:

  • saturation_curve.nc
  • saturation_curve_summary.csv
  • saturation_curve.png
  • forward_pass_contribution_curve.nc
  • forward_pass_contribution_curve_summary.csv
  • forward_pass_contribution_curve.png
  • adstock_curve.nc
  • adstock_curve_summary.csv
  • adstock_curve.png

These artefacts are intentionally different:

  • saturation_curve.* is the sampled saturation transformation on the scaled channel axis, exported with original-scale contribution values for easier reading. The PNG overlays that saturation-only curve against posterior mean realised contributions.
  • forward_pass_contribution_curve.* is a full-model direct contribution artefact. It rescales the observed historical spend path from 0% to 200%, runs that spend through the fitted adstock and saturation path, and records the resulting total channel contribution in original target units.
  • adstock_curve.* is the sampled carryover-weight profile for one impulse.

70_optimisation

This directory is present for every run, but the stage is skipped unless the YAML config contains an optimization block.

Main files when the stage runs:

  • optimized_allocation.nc
  • optimized_allocation.csv
  • response_distribution.nc (optimised expected media contribution)
  • reference_response_distribution.nc (reference allocation replay)
  • contribution_difference.nc (paired optimised minus reference draws)
  • budget_contribution_difference.csv (total paired difference summary)
  • optimize_result.json
  • budget_summary.csv
  • budget_response_points.csv
  • budget_impact.csv
  • budget_bounds_audit.csv
  • budget_roi_cpa.csv
  • budget_response_curves.csv
  • budget_mroi.csv
  • budget_optimisation.json
  • several PNG plots for allocation, contribution over time, response curves, impact, bounds audit, and ROI or CPA

The two response datasets use identical posterior draws and outcome dates, zero initial history, a shared time profile and no allocation or observation noise. current means a reference allocation replay, not fitted historical attribution. budget_optimisation.json records the comparison contract. budget_impact.csv includes channel-level paired difference HDIs; total difference HDIs come from summed paired draws, not sums of interval bounds.

Pipeline hdi_* fields now use ArviZ single-interval HDIs. Earlier pipeline exports used equal-tailed bounds under these names. Recompute summaries from retained draws; do not relabel historical files as corrected HDIs. Predictive coverage remains a separate equal-tailed calculation.

80_interpretation

Main files:

  • evidence_inventory.md
  • interpretation_report.md

Both files currently contain the same evidence inventory: preceding stage statuses, the holdout validation status and a reminder to review the retained artefacts. They do not provide automated business conclusions. A completed stage records successful report generation, not statistical qualification.

run_manifest.json

The manifest is the machine-readable index for the whole run.

Top-level fields include:

Field Meaning
run_name Effective run name
timestamp UTC run timestamp
config_path Original config path
output_dir Run directory path
status Overall run status
model_class Identifies the model object prepared in Stage 00; does not certify graph completion
data Basic dataset metadata
stages Per-stage manifest records
warnings Run-level warnings
error Run-level failure payload when the pipeline aborts

data includes:

  • x_shape
  • y_length
  • target_column
  • x_columns

Stage records

Each stage record contains:

Field Meaning
directory Stage directory name
status Current stage status
started_at ISO timestamp when the stage started
finished_at ISO timestamp when the stage finished
artifacts Mapping of artefact labels to root-relative paths
warnings Stage warnings
error Error string when the stage fails

The artifacts mapping uses root-relative paths such as 20_model_fit/model.nc.

Stage statuses

Status Meaning
pending Stage has not started yet
running Stage is currently running
completed Stage finished successfully
skipped Stage returned None intentionally
failed Stage raised an exception
not_reached A previous stage failed before this one ran

Common cases:

  • Stage 35 is skipped when validation is missing or disabled from YAML.
  • Stage 70 is skipped when optimization is missing from YAML.
  • Later stages become not_reached after the first failure.

Practical use

Use the run directory when you want:

  • a stable folder for downstream reporting
  • a machine-readable audit trail through run_manifest.json
  • stage-level links to artefacts without hard-coding filenames

If you want to add new artefact types or stages, see Extending the Runner.

CLI Reference

The pipeline exposes a thin CLI through abacus.pipeline.runner.

Entry point

python -m abacus.pipeline.runner --config path/to/config.yml

AI advisor patch approval uses a separate file-based entry point:

python -m abacus.pipeline.approval \
  --approval-request results/<run>/08_ai_advisor/approval_request.yaml

On success, the CLI prints the final run directory:

Structured pipeline completed: results/my_run_20260308_153000_a8k2m7q1

Arguments

Flag Required Default Meaning
--config Yes None YAML config path
--output-dir No results Root directory for pipeline runs
--run-name No Config filename stem Optional run-name override
--dataset-path No None Combined dataset CSV override
--x-path No None Feature CSV override when not using --dataset-path
--y-path No None Target CSV override when not using --dataset-path
--holidays-path No None Holiday CSV override
--target-column No None Target column used when reading CSV input
--prior-samples No 20 Prior predictive samples for Stage 10
--draws No None Posterior draws override
--tune No None Posterior tuning steps override
--chains No None Posterior chains override
--cores No None Posterior cores override
--random-seed No 42 Shared random seed
--curve-samples No 100 Posterior samples for Stage 60 curves
--curve-points No 100 Number of x-values for saturation curves

Common command patterns

Use the dataset path from YAML

python -m abacus.pipeline.runner \
  --config data/demo/geo_panel/config.yml

Override the combined dataset path

python -m abacus.pipeline.runner \
  --config configs/geo_panel.yml \
  --dataset-path /data/geo_panel_latest.csv \
  --run-name geo_panel_latest

Use separate feature and target files

python -m abacus.pipeline.runner \
  --config configs/panel.yml \
  --x-path /data/X.csv \
  --y-path /data/y.csv \
  --target-column revenue

Override sampler settings for one run

python -m abacus.pipeline.runner \
  --config configs/panel.yml \
  --draws 1000 \
  --tune 1000 \
  --chains 4 \
  --cores 4 \
  --random-seed 42

Override the holiday CSV

python -m abacus.pipeline.runner \
  --config configs/panel.yml \
  --holidays-path /data/holidays_uk_fr.csv

Approve an AI advisor patch

When ai_advisor.enabled: true and the advisor proposes a valid patch, Stage 08 writes 08_ai_advisor/approval_request.yaml with status: pending. Review the proposal, edit the request to status: approved, then run:

python -m abacus.pipeline.approval \
  --approval-request results/<run>/08_ai_advisor/approval_request.yaml \
  --approved-by "model owner"

The command writes approved_config.resolved.yaml and approval_record.yaml beside the advisor artifacts. It does not edit the source config.

How CLI overrides interact with YAML

The CLI does not replace the full YAML config. It only overrides the runtime fields exposed through PipelineRunConfig.

Important behaviours:

  • --dataset-path takes precedence over data.dataset_path.
  • --x-path and --y-path take precedence over data.x_path and data.y_path.
  • --holidays-path takes precedence over holidays.path.
  • --draws, --tune, --chains, --cores, and --random-seed are merged onto YAML fit.
  • --target-column affects CSV loading. Keep it consistent with target.column in YAML.

Exit behaviour

The CLI exits with status 0 on success. On failure, the process exits non-zero with the underlying exception.

The pipeline stops at the first stage failure. It does not provide flags to:

  • run only a subset of stages
  • continue after a failed stage
  • disable individual built-in stages other than omitting the optional optimization block from YAML

See Runner Overview and YAML Configuration for the execution model and config surface.

Extending the Runner

The retained runner is static, not plugin-based. To add a stage or integrate custom status reporting, extend the existing runner surfaces instead of bypassing them.

Stage contract

A stage function has this contract:

def run_some_stage(context: PipelineContext) -> dict[str, str] | None:
    ...

Return values:

  • return a dict[str, str] of artefact labels to root-relative paths when the stage succeeds
  • return None when the stage is intentionally skipped
  • raise an exception when the stage fails and should abort the run

The runner handles manifest updates around the stage call. Do not update context.manifest directly from a normal stage implementation unless you are changing core runner behaviour.

What is available in PipelineContext

PipelineContext gives each stage access to:

Field Use it for
run_config Runtime settings such as output root, seeds, and curve sample counts
raw_cfg The loaded YAML config as a mutable mapping
X, y Loaded dataset inputs
paths Stage directories and manifest path
manifest Current run manifest
model_kwargs Effective sampler overrides passed into model build
model Built PanelMMM, available after Stage 00

Artifact helpers

Use the helpers in abacus/pipeline/artifacts.py:

  • write_json(...)
  • write_dataframe(...)
  • write_dataset(...)
  • write_idata(...)
  • write_text(...)
  • save_figure(...)
  • copy_file(...)

Use context.paths.relative(path) when building the artefact mapping that the stage returns. The manifest expects root-relative paths, not absolute paths.

Adding a new stage

To add a new built-in stage, update these places:

  1. abacus/pipeline/artifacts.py Add the stage directory name to STAGE_DIRECTORIES.
  2. abacus/pipeline/runner.py Add a PipelineStageSpec to PIPELINE_STAGE_SPECS.
  3. abacus/pipeline/runner.py Add the stage function to the stage_functions mapping inside run_pipeline(...).
  4. abacus/pipeline/stages/__init__.py Export the new stage helper if you want it available from the stage package.

Minimal stage example

from abacus.pipeline.artifacts import write_dataframe


def run_custom_stage(context):
    if context.model is None:
        raise ValueError("Model has not been initialized before the custom stage.")

    stage_dir = context.paths.stage_dirs["custom"]
    output_path = stage_dir / "custom_summary.csv"

    frame = context.model.summary.total_contribution(output_format="pandas")
    write_dataframe(output_path, frame)

    return {
        "custom_summary": context.paths.relative(output_path),
    }

Optional stage pattern

If a stage should only run when a config block is present, follow the same pattern as Stage 70:

def run_optional_stage(context):
    cfg = context.raw_cfg.get("my_optional_block")
    if cfg is None:
        return None
    ...

Returning None is what marks the stage as skipped in the manifest.

Failure semantics

If your stage raises an exception:

  • the stage is marked failed
  • the run is marked failed
  • later pending stages are marked not_reached
  • run_pipeline(...) re-raises the exception

That means stage code should only catch exceptions when it can recover locally and still produce a valid artefact set.

Adding structured reporting

If you want progress callbacks without changing the core stage code, implement a PipelineReporter and pass it to run_pipeline(...).

The reporter protocol methods are:

  • on_pipeline_start(...)
  • on_stage_start(...)
  • on_stage_end(...)
  • on_pipeline_end(...)
  • on_pipeline_error(...)

This is the right extension point for:

  • notebooks or dashboards that want progress updates
  • lightweight orchestration wrappers
  • structured logging around pipeline runs

Consuming the manifest programmatically

The manifest is written after every stage transition, so external tools can poll run_manifest.json during execution. Each update replaces the file atomically from a completed temporary file in the same directory. Open the manifest again for each poll to observe later snapshots; a previously opened file can still refer to the older snapshot. Before initial publication the manifest may be absent. If publication fails, the last complete snapshot remains and the runner raises an error. A forcibly terminated process may leave a temporary file, which is not a published manifest.

The runner owns manifest updates during execution. Atomic replacement does not merge concurrent edits from other writers or make stage artefact writes transactional.

Typical uses:

  • check whether the optimisation stage was skipped
  • discover stage artefact paths without hard-coding filenames
  • detect the first failed stage and its error message

See Output Directory Schema for the manifest fields and status values.

Blocked Holdout Validation

Stage 35 is Abacus’s out-of-sample time-series validation layer.

It answers a narrower and more useful question than “does the model fit the training data?”:

“If I refit the MMM on the earlier history only, can it still predict the last blocked window reasonably well?”

For weekly MMM, that is usually a better stress test than a random split because media carryover, seasonality, and trend all depend on time order.

Why Abacus uses a blocked tail holdout

Abacus reserves the final holdout_observations unique dates as the holdout window, then fits a fresh model on the earlier dates only. The validation pipeline does not reuse Stage 20 posterior state.

That means Stage 35 is checking whether the full model specification can generalise forward in time, not whether the same fitted posterior can explain the rows it already saw.

This is especially useful in MMM because:

  • adstock depends on lagged spend history
  • seasonality is time-ordered rather than exchangeable
  • marketing calendars often drift near the end of the sample
  • overfit specifications can look fine in-sample and fail on the final weeks

See the implementation in validation.py.

How Stage 35 works

Given a YAML block such as:

validation:
  enabled: true
  holdout_observations: 8
  include_last_observations: true
  coverage_levels: [0.5, 0.8, 0.94]
  sampler:
    draws: 500
    tune: 500
    chains: 2
    cores: 2
    random_seed: 42

Abacus does the following:

  1. Sort all unique model dates.
  2. Reserve the last holdout_observations dates as the holdout window.
  3. Fit a fresh model on the remaining earlier dates only.
  4. Sample posterior predictive draws for the holdout rows.
  5. Compute uncertainty-aware predictive metrics and residual diagnostics.

If include_last_observations: true, Abacus prepends the trailing lag history needed for adstock carryover internally, then trims those prepended rows back out of the returned holdout predictions. This matters whenever media effects have memory.

Random seeds

Holdout posterior prediction explicitly receives the effective validation fitting seed. Seed precedence is the same as for the refit: base YAML fit, then runner sampler overrides, then validation.sampler. The metadata records prediction_random_seed alongside the effective sampler_config. A zero seed is valid; a missing or null effective seed leaves prediction unseeded and is recorded as null.

Reusing a seed supports repeatability within the same environment and execution settings; it does not promise identical draws across dependency versions or backends. No independent-stream derivation is applied.

Why this stage exists when Stage 30 already exists

Stage 30 and Stage 35 answer different questions.

  • Stage 30 is an in-sample fit check. It uses the same rows the model was fit on.
  • Stage 35 is an out-of-sample blocked holdout check. It uses future dates the validation fit did not see.

If Stage 30 looks good and Stage 35 looks weak, that is a classic warning sign of overfit, misspecification, or regime change.

Main artefacts

Stage 35 writes the following files under results/<run_name>_<timestamp>_<random_suffix>/35_holdout_validation/:

  • validation_metadata.json
  • holdout_posterior_predictive.nc
  • holdout_predictive_summary.csv
  • holdout_predictive_report.json
  • holdout_observed.csv
  • holdout_fitted.csv
  • holdout_residuals.csv
  • holdout_timeseries.png
  • holdout_residuals_acf.png

See also Output Directory Schema.

How to interpret the main metrics

The headline table is holdout_predictive_summary.csv.

The canonical predictive metric definitions specify the formulas, observation aggregation and missing-value boundaries.

Point-error metrics

RMSE gives larger errors more weight than MAE. NRMSE and NMAE divide those scores by the observed target range in the scored holdout, returning NaN when that range is approximately zero. Compare scores on comparable windows; range normalisation alone does not make different datasets comparable.

Use these as forecast-quality metrics, not causal-identification metrics.

Bias

bias is observed minus posterior predictive mean, averaged over observations:

  • positive bias means underprediction on average
  • negative bias means overprediction on average

For example, observed 12 and predicted mean 10 gives bias +2. Inspect persistent bias alongside trend changes, omitted predictors and changes in the holdout period. Its sign alone does not identify the cause.

CRPS

The continuous ranked probability score (CRPS) assesses the full predictive distribution. Lower is better on the same evaluation set. Inspect it alongside point errors, coverage and interval width; one score does not establish calibration.

Coverage

Stage 35 reports coverage_50, coverage_80 and coverage_94, the observed fractions inside equal-tailed posterior predictive intervals at those nominal probabilities. They do not measure coverage of mmm.summary HDIs. See the metric definitions for the quantiles, endpoint inclusion and finite-observation denominator.

Compare coverage with its nominal probability alongside bias, interval width, residual patterns, holdout size and dependence between observations. Low coverage can reflect narrow intervals, prediction bias or distributional change. High coverage does not by itself prove that intervals are too wide.

With eight finite aggregate observations, coverage changes in steps of 1/8. Seven covered observations give 0.875, so exact agreement with 0.94 is impossible. That result alone cannot establish or refute nominal 94% calibration. Serial or panel dependence further limits the information in a short holdout; panel rows do not automatically supply independent evidence.

How to read the plots

holdout_timeseries.png

This is the first plot to inspect.

Look for:

  • whether the observed series generally stays inside the predictive intervals
  • whether misses are isolated or systematically one-sided
  • whether the model misses turning points or holiday spikes
  • whether the predictive band width looks plausible relative to the volatility of the target

Common interpretations:

  • repeated misses on the same side: likely bias
  • clustered misses: inspect bias, changing volatility and omitted time structure
  • wide bands: inspect predictive spread and its sensitivity to the specification; width alone does not diagnose identification

holdout_residuals_acf.png

This checks whether the holdout residuals still contain serial structure.

Look for:

  • obvious positive autocorrelation across nearby lags
  • repeating seasonal patterns
  • long runs of same-sign residuals

If residual autocorrelation is strong, the model is usually still missing some time structure such as:

  • seasonality
  • holiday dynamics
  • delayed media effects
  • structural breaks or time-varying baseline behavior

Practical rules of thumb

These are pragmatic MMM heuristics, not hard pass/fail thresholds.

Choosing the holdout window

  • For weekly MMM with roughly 1 to 2 years of data, 6 to 12 weeks is a practical starting range.
  • 8 weeks is a sensible default when you want enough tail signal without throwing away too much history.
  • If the dataset is very short, a larger holdout can make the validation noisy and can leave too little training history for stable estimation.

Abacus itself enforces that blocked holdout validation must still leave enough training dates for the model to run; see validation.py.

Comparing Stage 30 and Stage 35

  • Expect Stage 35 to be worse than Stage 30. That is normal.
  • Worry when the degradation is large or the direction changes materially.
  • If Stage 30 is excellent and Stage 35 is weak, suspect overfit or misspecification before celebrating the in-sample fit.

Reading coverage

  • Report nominal probability, empirical coverage and the number of scored observations.
  • Investigate departures alongside bias and width; do not infer the cause from coverage alone.
  • Use additional comparable windows when feasible; one short holdout does not establish calibration.

Reading bias

  • A small nonzero bias is normal.
  • Large consistent bias over the holdout tail is a stronger warning sign than a single noisy miss.
  • If bias keeps the same sign across several model variants, inspect trend, holidays, and baseline structure before changing media priors.

Reading residual structure

  • White-noise-like residuals are what you want.
  • Visible residual runs or autocorrelation usually mean the model is still missing systematic time variation.
  • Do not treat a decent RMSE as sufficient if residuals still show structure.

What Stage 35 does not tell you

Blocked holdout validation is valuable, but it is not a causal guarantee.

It does not prove:

  • that channel attribution is identified
  • that ROAS is unbiased
  • that the chosen priors are correct
  • that the model is safe for large budget reallocation on its own

It does tell you whether the specification can forecast a held-out tail window coherently. That makes it an important diagnostic, but still only one part of MMM model assessment.

For most weekly MMM work:

  1. Run Stage 30 and Stage 35 together.
  2. Compare in-sample and holdout metrics before changing the specification.
  3. Use the same holdout window across candidate models so the comparison is fair.
  4. Prefer specifications that are stable across reasonable prior choices, not just the one that scores best on a single holdout.
  5. Treat Stage 35 as a forecasting sanity check alongside prior predictive checks, posterior predictive checks, and substantive business review.

Common mistakes

  • Using a random split instead of a blocked time split
  • Reading Stage 30 as out-of-sample validation
  • Ignoring coverage and focusing only on RMSE
  • Using a holdout that is too long for the amount of history available
  • Repeatedly tuning the spec to one holdout window until the score looks good