FAQ

This section collects longer-form answers to recurring MMM, Bayesian, and panel-econometrics questions that come up when practitioners move from classical econometrics to PanelMMM.

The pages are written for technical readers who already understand regression, panel data, and causal inference, but want the Abacus framing. Start with the practical questions to connect those concepts to an analyst’s workflow.

Practical analyst questions

Core model concepts

Priors and model checking

Computation and comparison

Panel specification

Suggested reading order

If you are new to Bayesian MMM, a practical sequence is:

  1. Data suitability and preparation
  2. Model choice and supported operations
  3. Causal Identification in Marketing Mix Modelling
  4. Bayesian Priors for Econometricians
  5. Prior Predictive Checks for Econometricians
  6. MCMC Diagnostics for Econometricians
  7. Posterior Predictive Checks for Econometricians
  8. Model Comparison for Econometricians
  9. Contributions, ROAS and scenarios

Subsections of FAQ

Is my data suitable for Abacus, and what must I prepare?

Check two things before fitting: whether the data satisfy Abacus’s input contract, and whether they contain enough relevant variation to answer your question. A correctly formatted dataset can still support weak or ambiguous channel estimates.

How much history do I need?

There is no universal minimum number of weeks, markets or observations that makes an MMM reliable. Assess the history against the proposed model:

  • Does it cover the seasonal patterns and business regimes you want to model?
  • Do channels change independently enough to separate their responses?
  • Is there information about different exposure levels, including the range relevant to the proposed decision?
  • Is there enough earlier history for the chosen adstock horizon and enough data left after reserving a meaningful holdout?
  • For a panel estimator, is there usable variation within units and, where required, between units?

Three years of channels that always move together may contain less useful channel-specific information than a shorter period with distinct changes. Adding markets does not automatically add independent information either. Priors can regularise weakly informed parameters; they do not create the missing variation. See Bayesian Priors.

Should my media inputs be spend, impressions or something else?

Use a consistently measured exposure variable that matches the response you want to model. Abacus applies adstock and saturation to the channel columns you supply. It does not obtain a separate monetary spend series or convert impressions into money for you.

The choice also determines which downstream interpretation is valid:

Channel input Outcome Contribution divided by input means
Monetary spend Revenue in the same currency Model-conditional ROAS
Monetary spend Conversion count Conversions per unit of spend; its reciprocal is cost per conversion
Impressions Revenue Revenue per impression, not ROAS

The standard budget workflow treats allocated channel-input quantities as spend. Do not pass a monetary budget into a model trained on impressions and assume Abacus supplies the price conversion. That requires an explicit, reviewed mapping and a compatible planning workflow.

Keep currency, tax treatment, gross/net definitions and attribution dates consistent. A change in measurement can resemble a change in response.

What table must I provide?

For Python fitting, supply X as a DataFrame and y as a row-aligned Series or one-dimensional array. Keep the target out of X; include the date, channel, control and declared panel-dimension columns. A combined CSV is supported by the YAML and runner workflows, which separate the target for you.

For an aggregate time series, use one observation per date. For a panel, provide exactly one row for each expected date and panel-coordinate combination. Multiple panel dimensions require the declared rectangular grid; they do not represent arbitrary nested or incomplete groups.

Use a consistent time frequency and align the outcome, media and controls to the same periods. l_max=8 refers to eight input periods, including the current period, not necessarily eight weeks. Document aggregation rules: sums can make sense for spend and revenue, while a price or rate needs a justified aggregation method.

See Input Data Requirements and Panel Data Layout for the exact column and alignment contracts.

Are missing values the same as zero activity?

No. Abacus requires observed values for the channel, control and target cells you supply; it does not silently replace missing measurements with zero. Absent panel rows and duplicate rows also need resolution before fitting.

For example, consider two markets in one week:

Market Recorded TV spend Interpretation and action
North 0 Keep zero if the source confirms no TV spend occurred
South Missing Investigate the missing measurement; do not infer zero spend

If South’s sales are also missing, adding a row with zero sales invents an observation. Recover the measurement or choose and document a defensible treatment of the missing data. Any imputation introduces assumptions whose effect on the analysis should be assessed. If you change the analysis window or units, check the panel contract again.

What does Abacus preprocess automatically?

Abacus scales the target and channel inputs before constructing the model. It does not automatically scale controls, choose a control set, reconcile currencies, adjust for inflation or perform domain-aware missing-data repair.

Default target and channel scaling pools over the panel dimensions. It is not automatically separate scaling for each market. Check the configured scaling dimensions before choosing prior magnitudes or comparing parameters.

Supply the media inputs intended for the configured transforms. Do not pre-adstock or pre-saturate them and then apply the same transformations again inside Abacus. Check the outcome implications of your complete specification with prior predictive draws on the appropriate scale.

See Scaling and Preprocessing and Prior Predictive Checks.

Should I combine correlated channels or add more controls?

Neither action is an automatic remedy. If two channels always run together, separate contributions may depend heavily on priors and specification choices. Consider a combined channel when it represents a coherent exposure and the decision can be made at that combined level. Combining channels changes the estimand: it cannot then justify separate channel allocations. Changing media mix or unit costs can also make the combined response unstable.

If separate estimates are essential, seek additional identifying variation or relevant experimental evidence, and report the limits of the current data. Inspect transformed predictors as well as raw correlations; the model learns from adstocked and saturated exposures.

Choose controls from their role in the outcome and media-assignment process, not solely because they improve fit. A variable affected by media can block part of the effect you intended to estimate. More controls can also introduce collinearity. See Causal Identification.

What should I prepare before the first fit?

Record the target, units, time frequency, panel definition and decision question. Retain an audited input table, its source and all preprocessing decisions. Review variation, plausible controls, model scale and priors. Choose the estimator and reserve the validation window before comparing preferred results.

Use Blocked Holdout Validation for a training-prefix refit and later holdout. The supplied holdout media and controls make this a conditional prediction exercise; it does not establish that those inputs would have been known in a live forecast. Keep information from the holdout out of fitted preprocessing and model selection wherever that information would be unavailable at prediction time.

Continue with Which Abacus model should I choose?.

Which Abacus model should I choose, and what does that choice permit?

Choose the model from your question, data structure and defensible assumptions before examining preferred channel results. In Abacus, the choice also determines which prediction, calibration and planning operations are available.

This page describes the accompanying Abacus 3.1.1 library. The estimator support matrix is the detailed reference for named presets.

Which model is the starting point for my data?

Situation Candidate What you must justify
One aggregate observation per date Named time_series preset Temporal variation, baseline, controls and media response specification
A rectangular panel with a deliberately configured parameter and prior structure Ordinary dimensioned PanelMMM Which parameters vary by slice, which are shared and whether any hierarchical pooling is specified
A balanced panel where shared slopes should use changes within units Named fe preset Sufficient transformed within-unit variation and the remaining time-varying confounding assumptions
A balanced panel requiring an explicit adjustment for persistent unit differences associated with predictors Named cre preset The declared transformed between-unit summary basis, its estimability and sufficient within-unit variation

The named FE and CRE presets take one unit dimension, such as geo, containing multiple units. This is not a claim that they fit only one market. Their released media slopes and adstock/saturation parameters are shared across units. Do not interpret a separate regional contribution as a separately estimated regional response function.

Does setting dims select FE or create partial pooling?

No. dims=("geo",) declares an array axis. On the ordinary PanelMMM surface, the parameter dimensions and priors determine which quantities are shared or vary by market. Independent market-indexed priors are not hierarchical partial pooling; shared parameters are not the named FE likelihood.

Specify a named estimator explicitly when you want its contract. The legacy use_mundlak_cre=True option on an ordinary panel model is also not an alias for the named CRE preset. See Panel Dimensions and the Mundlak FAQ.

How do FE and CRE differ in the information they use?

FE removes persistent unit intercepts from its shared-slope likelihood by using within-unit contrasts. It cannot learn a slope from a predictor that does not vary within any unit. A larger market’s consistently higher spend and sales do not supply the same information as changes within that market.

CRE models unit intercepts together with declared centred between-unit summaries. In the named Abacus preset, media summaries use fitted transformed exposures, not simply raw-spend averages. The adjustment is limited to that declared basis and needs usable between-unit information as well as within-unit variation.

For example, suppose larger markets always spend more and sell more. FE asks what the within-market changes reveal under its model. CRE additionally represents the specified relationship between persistent market differences and predictor summaries. If spend barely changes within markets, neither preset creates the missing within-market information.

Neither removes arbitrary time-varying confounding. The released named presets also do not support common categorical time effects. Read the FE specification and CRE specification before importing assumptions from another panel package.

Why is the declared re preset unavailable?

Implementing and qualifying the named random-effects preset was a lower priority because its assumptions are often difficult to defend in marketing applications.

Standard RE requires the unobserved persistent unit effect to have conditional mean zero given the included predictor history. Marketing budgets commonly reflect market size and expected baseline demand, which also affect sales. When the model does not adequately account for these factors, the RE restriction is implausible.

Correlation between spend and observed market size is not itself a violation if the specification adequately accounts for market size. The restriction concerns the remaining unobserved unit effect; it is not a rule that raw spend and market size must be uncorrelated.

Abacus recognises estimator.type: re in configuration, but the named preset has not passed the required implementation checkpoint for public release. Building it raises EstimatorReleaseGateError before fitting. Internal implementation work does not make it a released estimator. This is an intentional restriction, not an installation problem or a general statistical objection to random-effects modelling.

Choose another estimator only when its assumptions fit the question. CRE is not an automatic substitute for every RE analysis.

Can every model predict, calibrate and optimise?

No. For the named presets in this release:

Operation time_series fe cre
Prediction on new dates with supplied inputs Supported All fitted units required All fitted units and frozen fitted CRE summaries required
Historical and manual allocation scenarios Supported Supported for all fitted units Supported with the same fitted-unit and frozen-summary restrictions
Lift-test or cost-per-target calibration Supported through the ordinary model surface Unavailable Unavailable
Fixed-budget optimisation Supported Unavailable Unavailable

Supported operations still require valid inputs and a suitable fitted model. For ordinary dimensioned PanelMMM, check the configured model and operation contracts rather than assuming the named-preset table describes every custom configuration. New dates do not imply support for previously unseen markets.

For panel manual allocations, use a labelled xarray.DataArray or DataArraySpec covering the required fitted coordinates. A channel-only dictionary is not a panel allocation. See Scenario Specifications.

An unsupported-operation error is not resolved by removing the estimator label from a fitted model. That would discard the contract without making the operation statistically valid.

Should I use Python, a YAML builder or the pipeline runner?

These are workflow choices rather than competing estimators. Python gives direct control over construction, fitting and analysis. The YAML builder constructs a model from configuration for subsequent use. The runner executes a staged workflow and writes artefacts to disk.

Their configuration surfaces differ. Start with the appropriate Python, YAML builder or runner guide. A successful build or run does not establish that the chosen model is suitable for the data.

Can I choose the model with the best LOO score?

Only compare scores that describe the same prediction task, observations, outcome scale and likelihood measure. Abacus’s named FE likelihood scores within-unit contrasts; CRE scores complete unit blocks with its unit intercept integrated out. Those scores cannot be ranked against each other as though they evaluated identical outcome-level observations.

Use the Model Comparison FAQ to check comparability. Assess computation, predictive adequacy, sensitivity and identification separately. Choose the model because its contract answers the question, then report what the evidence permits.

What do contribution, ROAS and scenario results actually mean?

Read each output as a quantity defined by the fitted model, its units and its evaluation window. A contribution table, revenue prediction and allocation comparison answer different questions. None independently establishes a causal media effect.

Is media contribution the same as predicted revenue?

No. In the ordinary additive model, media contribution is the media component of the fitted mean. The full mean also includes the configured intercept, controls, seasonality and other additive terms. Posterior predictive outcome draws additionally include observation uncertainty through the likelihood.

Output What it describes What it does not automatically describe
Historical channel contribution A component evaluated for the observed media path under the fitted model A measured causal increment
Expected media contribution under a scenario The model’s media response to the specified input path Total future revenue or realised sales
Posterior predictive outcome An outcome draw conditional on supplied inputs and the fitted model Uncertainty about every possible future input or structural change
Optimised allocation A solution for the configured objective, budget and constraints An approved business recommendation

Named FE and CRE have their own likelihood and prediction contracts; do not transfer every ordinary-model interpretation without checking those contracts. See Model Overview and Contributions and Decomposition.

When is a reported ratio really ROAS or CPA?

The efficiency accessors use fitted contributions and the supplied channel inputs. They do not look up an independent spend series.

  • Revenue contribution divided by monetary spend in the same currency gives model-conditional return on advertising spend (ROAS).
  • Monetary spend divided by conversion contribution gives model-conditional cost per acquisition/conversion (CPA), with the conversion definition stated.
  • Revenue contribution divided by impressions is revenue per impression, even if an output column is labelled ROAS.

For example, £200 of modelled revenue contribution divided by £100 of spend gives ROAS 2. £200 divided by 1,000 impressions gives £0.20 per impression. Neither calculation is profit: margins and other costs are separate.

Use original-scale contributions and consistent periods and panel aggregation. target_type selects an efficiency accessor and label; it does not validate currencies or convert exposure units. The element-wise accessors return NaN for a zero denominator. Do not replace that undefined ratio with a favourable return. See ROAS and Metrics.

How should I aggregate returns and uncertainty?

Define the aggregate quantity first. For a total-window return, sum contributions within each posterior draw over the intended dates and regions, then divide by the corresponding total spend. Summarise the resulting draw distribution. An unweighted average of weekly or regional ratios generally answers a different question.

For example, spend of £100 and £900 with contributions of £300 and £900 gives individual ROAS values of 3 and 1. Their unweighted mean is 2, but the combined return is £1,200 / £1,000 = 1.2.

Similarly, sum contributions within each draw before calculating an interval for the total. Adding component interval endpoints does not generally give the total’s credible interval because the components are dependent. Inspect the aggregation behaviour of the specific accessor you use; a frequency argument alone does not define the desired ratio estimand. See Summary and Export.

Why might optimisation favour a channel with lower historical ROAS?

Historical average ROAS describes return over the evaluated exposure path. Allocation decisions depend on the change in the objective from feasible changes in spend. With saturation, a channel can have high historical average return but low marginal response at its current spending level.

For the low-level PanelBudgetOptimizerWrapper, the default objective is average posterior total_media_contribution_original_scale. It is not automatically profit, a lower credible bound or a risk-adjusted business utility. The feasible solution also depends on bounds, time allocation and carryover assumptions.

A channel at its upper bound may indicate that the objective would prefer more spend if allowed. It does not establish that the boundary is a validated commercial optimum. Inspect the solver result, constraint satisfaction, supported spending range and sensitivity before interpreting it. See Budget Optimisation.

Is the budget per period or for the whole window?

Check the entry point; the contracts differ.

Entry point Budget or allocation units
PanelBudgetOptimizerWrapper.optimize_budget(budget=...) Total across allocation cells for one model period
ManualAllocationScenarioSpec.allocation Total over the requested spending window for each allocation cell
FixedBudgetOptimizedScenarioSpec.total_budget Total over the requested spending window
Preferred runner optimization.budget block Total-window budget, resolved according to its configured mode
Legacy runner optimization.total_budget Per-period budget

For eight weekly spending periods, a flat £80,000 window budget corresponds to £10,000 per period across all allocation cells. Passing budget=80_000 to the low-level wrapper instead requests £640,000 over those eight periods.

Panel allocations also need the correct region/channel coordinates. See Scenario Specifications and YAML Configuration.

Why do historical and simulated scenarios give different results?

A historical reference uses observed history. A simulated allocation defines a spending path, which may differ in timing even when its channel totals match that history. Adstock and saturation make timing consequential.

Check the requested spending window, evaluated response window, incoming lag history, time distribution and carryover tail. With include_carryover=True, the response window extends beyond the spending window. A default simulated scenario uses include_last_observations=False, so historical carryover is not automatically included.

noise_level controls simulated spend-path variation; it is not the outcome likelihood’s residual uncertainty. Set it to zero when you need a deterministic spend path. Read the resulting metadata rather than inferring the scenario definition from its name.

CurrentScenarioSpec requires overlap with observed dates. It is not a future-only no-change forecast. See Overview and Workflow.

How do I compare two plans and report uncertainty in their difference?

Define a common response quantity and align the plans’ response horizon, initial history, carryover treatment and posterior draws. For a fixed-budget reallocation question, keep total spend the same. An expansion plus reallocation answers a different question and should be labelled accordingly.

The following is illustrative arithmetic, not a fitted Abacus result:

Plan Eight-week spend Mean media contribution over the same response window
Reference allocation replay £80,000 £160,000
Alternative allocation replay £80,000 £176,000

The mean difference is £16,000. That table alone supplies no credible interval for the difference. Compute the alternative minus reference contribution for each matched posterior draw, then summarise those differences. Do not subtract the endpoints of the two marginal intervals or infer the difference’s uncertainty from whether those intervals overlap.

If the alternative instead spends £88,000, report an expansion-plus-reallocation contrast. If the reference is historical attribution with different incoming history, re-evaluate a matching reference path before making the controlled allocation comparison.

ScenarioPlanner.compare(...) concatenates individual scenario summaries; it does not by itself provide a paired difference interval. The pipeline’s matched allocation comparison retains separate paired-difference artefacts. Check the actual output contract in Comparison Outputs and Output Directory Schema.

What is required before recommending a budget change?

Review computation, prior plausibility, predictive checks and relevant holdout evidence. Then assess identification, parameter and specification sensitivity, extrapolation, commercial constraints and the uncertainty in the decision quantity. A solver success flag or narrow posterior interval does not replace these checks.

State the response quantity, units, spending and response windows, assumptions and permitted interpretation with the recommendation. Posterior uncertainty conditions on the fitted model; it does not automatically include omitted confounding, model-choice uncertainty or future structural change. See Causal Identification and Interpreting Optimisation.

Bayesian Priors

Priors constrain and regularise the model. Their influence depends on their support and scale, the likelihood and the information in the data. A proper posterior or stable fit does not by itself establish identification.

Support restrictions and prior scale

Separate two choices:

  • Support: which parameter values the model permits.
  • Concentration: how it distributes probability over those values.

A HalfNormal prior gives zero probability to negative values. A LogNormal prior restricts its parameter to strictly positive values. Increasing either prior’s scale does not allow the posterior to become negative: the likelihood cannot create posterior support where the prior assigns none.

Consequently, P(beta > 0 | data) is one by construction for a coefficient with a positive continuous prior. It is not evidence that the data established a positive effect. An interval above zero must be interpreted in light of that restriction. A posterior can still concentrate near zero; whether it rules out a practically negligible effect is a separate question.

A concentrated prior can reduce variance while introducing bias when its assumptions are wrong. A diffuse prior can permit implausible values or weakly identified combinations. Neither choice guarantees accurate estimates, good sampling or causal validity. “Weakly informative” only has meaning relative to the parameterisation and input/output scales.

Relation to penalised estimation

Under a specified likelihood, a Normal prior yields a quadratic penalty in the negative log posterior, and a Laplace prior yields an absolute-value penalty. This connects posterior modes to ridge and lasso estimates in the corresponding models. It does not make their full uncertainty calculations identical.

An unconstrained frequentist estimate does not require a Bayesian prior. A flat density on the whole real line is improper, and flatness changes under non-linear reparameterisation. Avoid treating it as an assumption-free probability distribution.

Specify priors in Abacus

Use Prior from pymc_extras.prior. For example, these transform objects restrict the response amplitude to be non-negative and the decay to (0, 1):

from pymc_extras.prior import Prior

from abacus.mmm import GeometricAdstock, LogisticSaturation

adstock = GeometricAdstock(
    l_max=4,
    priors={"alpha": Prior("Beta", alpha=1, beta=3)},
)
saturation = LogisticSaturation(
    priors={
        "beta": Prior("HalfNormal", sigma=2),
        "lam": Prior("Gamma", alpha=3, beta=1),
    },
)

This is a configuration fragment: pass the objects as adstock and saturation when constructing PanelMMM. The numerical scales above are illustrative, not a recommendation for every dataset. Abacus applies media transforms to scaled inputs, so assess their implications on the outcome scale too. For model-level overrides and valid YAML syntax, use Priors and Configuration.

Choose and assess a prior specification

  1. State the support restrictions and their substantive justification.
  2. Check plausible parameter magnitudes in the actual model scales.
  3. Inspect prior predictive draws for plausible outcomes before interpreting a fit.
  4. Assess the available identifying variation, including correlated media, persistent unit differences and possible time-varying confounding.
  5. Fit defensible alternative prior specifications and compare the quantities used for decisions, including contributions and scenario contrasts.

There is no universal threshold in weeks or number of channels that makes media effects identified or a prior negligible. More observations need not supply independent variation. External calibration can inform a particular response, but its relevance depends on the experimental design, estimand and transport assumptions; see Calibration.

Compare like quantities

Compare each parameter prior with its parameter posterior to inspect updating. Similar distributions do not uniquely diagnose an excessively concentrated prior: the likelihood may be weak, compatible with the prior, or informative about combinations rather than individual parameters. A shifted posterior also does not show that the prior has ceased to matter.

Compare prior predictive and posterior predictive distributions with observed outcomes to assess their implications for data. These distributions include different sources of uncertainty from a parameter distribution and cannot be substituted for it. Sensitivity analysis is needed to assess how conclusions depend on the prior, even when the posterior looks concentrated.

For the broader distinction between fit and attribution, see Baseline vs Media Trade-Offs and Causal Identification.

Adstock and Saturation

Adstock represents carryover; saturation represents a non-linear response to media exposure. In PanelMMM, the selected transform families and their priors define the response model. Joint estimation propagates uncertainty in those parameters, conditional on the specified model; it does not remove the need to assess identification and functional-form sensitivity.

Geometric adstock and its finite horizon

For the usual lagged convolution, with L = l_max, geometric adstock uses weights proportional to $\alpha^\ell$ for $\ell=0,\ldots,L-1$. With the GeometricAdstock object’s default normalize=True, the transformed series is

$$ x_t^* = \frac{\sum_{\ell=0}^{L-1}\alpha^\ell x_{t-\ell}} {\sum_{\ell=0}^{L-1}\alpha^\ell}. $$

With normalize=False, omit the denominator. This is a finite convolution, not an untruncated recursive Koyck model. l_max=4 retains the current period and three preceding periods. The available history and convolution mode also matter, especially at the start of a series or prediction window.

The default decay prior is Beta(alpha=1, beta=3). A decay near zero gives little weight to older exposure; a decay near one retains substantial weight across the specified horizon. The prior does not ensure that omitted lags are negligible. Choose l_max using plausible carryover and assess sensitivity to that choice. The lag count refers to the input frequency, not necessarily weeks.

Alternative transform families encode different lag-weight shapes. Inspect their parameterisation rather than assuming that every alternative permits a delayed peak. See Adstock and Saturation for configuration and supported classes.

The retained logistic saturation function

LogisticSaturation applies

$$ f(x) = \beta\frac{1-e^{-\lambda x}}{1+e^{-\lambda x}} = \beta\tanh(\lambda x/2). $$

For positive beta and lam on non-negative inputs, the response is concave and approaches beta. It reaches half that limit at $x=\log(3)/\lambda$. Increasing lam moves half-saturation towards zero; it does not create a freely located positive-spend inflection point. This is not a general sigmoid with a learnable threshold before diminishing returns begin.

The default priors are Gamma(alpha=3, beta=1) for lam and HalfNormal(sigma=2) for beta. Their implications depend on the model’s scaling: the transform receives scaled inputs, and beta sets the response amplitude on the model scale. See Scaling and Preprocessing and Bayesian Priors. Positive support is a modelling restriction, not evidence learned about the sign.

Joint estimation and omitted uncertainty

Abacus estimates the uncertain transform parameters jointly with the other model parameters. In this logistic specification, beta is the response amplitude; it is not an additional coefficient to be multiplied by a separate generic media slope.

Fixing transform parameters before fitting conditions on those choices. Joint estimation can propagate their posterior uncertainty, but still conditions on the selected transform families, lag horizon, controls and baseline. Neither approach guarantees narrower, wider or better-calibrated intervals in every dataset. Assess sensitivity in the quantities used for decisions, not only in parameter summaries.

Transformation order

adstock_first=True, the PanelMMM default, applies saturation to accumulated media exposure. adstock_first=False applies saturation within each period before carrying the response across periods. These operations generally do not commute.

Choose the order from the response mechanism and assess its implications; channel names alone do not determine it. Normalisation and lag length affect the result as well. A good fit under either order does not prove that its media decomposition is correct. See Baseline vs Media Trade-Offs.

HSGP

A Hilbert space Gaussian process (HSGP) approximates a Gaussian process with a finite set of basis functions. In Abacus it provides a regularised function for components such as a smooth baseline. Basis size, covariance assumptions and hyperpriors all affect what that component can represent.

Basis size and model flexibility

The m setting controls the number of retained basis functions. Conditional on the GP hyperparameters, Abacus gives the basis coefficients Normal priors whose scales depend on the covariance’s spectral density. These priors can shrink high-frequency terms; they do not generally set them exactly to zero.

The number of basis coefficients is not an OLS residual-degrees-of-freedom calculation. Regularisation can make effective flexibility smaller than the basis count, but the data do not automatically select a uniquely correct amount of flexibility. Poorly constrained hyperparameters and baseline/media trade-offs can remain even when fitting succeeds.

Select and check the approximation

HSGP.parameterize_from_data(...) recommends m and the boundary extent from the supplied data and lengthscale settings. It provides a starting point, not proof that the approximation is adequate for the fitted posterior.

A basis that is too small can miss variation permitted by the covariance. Increasing m can change the fitted function until the approximation is adequate, and increases computation. There is no guarantee that m=50 and m=500 produce identical curves. Assess both approximation settings and hyperprior sensitivity; do not choose them solely to obtain a preferred media result. Riutort-Mayol et al. describe basis and boundary selection and diagnostics for approximation adequacy.

Baseline and media remain competing explanations

A regularised baseline can still explain variation that also aligns with media. Orthogonality among mathematical basis functions does not imply orthogonality to the observed media design. Shrinkage changes the allocation of variation; it does not remove omitted confounding or identify the causal media contribution.

Check whether substantive results change across plausible baseline and prior specifications, as well as whether the total fit changes. See Baseline vs Media Trade-Offs.

Choose between Fourier and HSGP seasonality

Component Assumption and practical consideration
YearlyFourier A finite seasonal basis with the configured coefficient priors; order controls retained harmonics
HSGP A finite approximation to a non-periodic GP; covariance and hyperpriors control smoothness and scale
HSGPPeriodic A finite periodic GP approximation; the seasonal function repeats at its configured period

The retained HSGPPeriodic does not introduce slowly drifting seasonal coefficients. Modelling changing seasonal shape requires an explicit model for that change; choosing this class alone does not provide one.

HSGP uses a basis representation rather than the full observation covariance factorisation of an exact GP. Its cost still depends on basis size, model structure and sampling behaviour. Neither HSGP nor Fourier is universally superior. Use Seasonality and Trends for supported configuration, then assess predictive behaviour and attribution sensitivity for the intended task.

Model dated events explicitly

A holiday indicator and a smooth dated-event basis encode different temporal shapes. Choose according to the event’s expected duration and the data; a smooth build-up and decay is not always more realistic than an indicator.

For event attachment, use the example and prerequisites in Additive Effects and Events. Attach the effect to an unbuilt model, then build and fit with matching X and y. Event effects add explanatory components; they are not lift-test calibration. Check that reference’s persistence limitation before saving a model containing events.

MCMC Diagnostics

Use MCMC diagnostics to assess whether retained simulation draws support the posterior summaries you want to report. They concern Monte Carlo exploration and precision. They do not establish that the model is correctly specified, that media effects are identified, or that predictions will generalise.

For Abacus methods, report fields, threshold comparisons and unavailable states, use the canonical diagnostics guide.

1. What the sampler does

Markov chain Monte Carlo (MCMC) approximates posterior expectations and probabilities using dependent simulation draws. Abacus’s usual NUTS workflow uses Hamiltonian Monte Carlo trajectories with an adaptive trajectory length. Warmup adapts sampling settings; retained draws are used for posterior summaries. A larger retained draw count does not by itself establish adequate exploration.

For example, two chains with 2,000 retained draws each provide 4,000 parameter vectors. Means, intervals and estimated posterior probabilities derived from those vectors have Monte Carlo error. Keep that simulation error distinct from posterior uncertainty about a parameter.

2. Trace and rank plots

Inspect multiple chains for drift, persistent differences, long periods of little movement and uneven exploration. Agreement across chains supports the numerical assessment, but chains can agree while missing the same region. Trace plots need not look like independent white noise: MCMC draws are dependent by construction.

Use trace or rank plots alongside numerical diagnostics, including checks of the parameters and derived quantities that drive the decision. A visually stable trace alone is insufficient evidence of convergence.

3. R-hat: screen for disagreement

Modern R-hat uses split chains, rank normalisation and a folded comparison to detect differences in location and scale. It improves on the original between-chain/within-chain variance comparison. See Vehtari et al. for the method and limitations.

Values close to one are desirable. Abacus’s default 1.01 threshold is a screening convention, not a safe/unsafe boundary that proves convergence. Values at or above the configured threshold are flagged. Non-finite or unavailable diagnostics must be investigated rather than counted as passing.

When chains disagree, inspect their paths and the model geometry. More warmup or retained draws may help, but persistent separation can require a different parameterisation or a revised model. Simply extending a run is not a general solution.

4. Effective sample size: Monte Carlo precision

Effective sample size (ESS) describes the precision of a simulation-based summary relative to independent draws. It is not the number of observations in the dataset or the model’s degrees of freedom. It is also not a Newey–West standard-error correction: that procedure concerns estimation uncertainty under dependent data, whereas MCMC ESS concerns dependent simulation draws.

Bulk ESS screens exploration of the main posterior mass; tail ESS helps assess precision near interval endpoints. Adequacy depends on the quantity and precision needed. Abacus’s default ESS threshold of 400 is a screening convention, not a universal guarantee. Check Monte Carlo standard errors for reported summaries; rare-event probabilities can need substantially more simulation than a posterior mean.

If exploration is otherwise adequate, more retained draws can improve precision. If chains mix poorly, investigate scaling and parameterisation. Thinning an existing chain discards draws; it does not recover unexplored regions and is not a general remedy for low ESS. Storage constraints are a separate consideration.

5. Divergences: investigate retained transitions

A divergence signals excessive numerical error along a Hamiltonian trajectory and can indicate posterior geometry that the sampler explores poorly. It is evidence of a computational problem to investigate, not proof of one particular model defect or a known amount of bias.

Separate warnings during warmup from divergences among retained draws. Target zero retained divergences and investigate any that remain, even when R-hat and ESS look acceptable. Do not dismiss a small retained count by calling it warmup. Locate the affected parameter regions and examine sensitivity to sampling settings and parameterisation. See the Stan diagnostic guidance for the computational interpretation.

Increasing target_accept can reduce integration error at additional computational cost. Persistent divergences can require rescaling, reparameterisation or revising the model. Adding draws alone does not repair the underlying geometry. Zero observed divergences is useful evidence, but not proof of adequate exploration or model validity.

Abacus distinguishes unavailable divergence evidence from an observed zero. Missing or invalid retained flags produce an unavailable status and reason; they must not be interpreted as a clean run. Consult the diagnostics guide for the actual report fields and pipeline handling.

6. Interpret posterior intervals conditionally

A 95% credible interval contains 95% of the parameter’s posterior probability, conditional on the data, likelihood and prior. That statement is different from the repeated-sampling coverage of a confidence-interval procedure. Neither interpretation removes the need to assess model assumptions.

A highest-density interval (HDI) and an equal-tailed interval are different summaries and can differ for skewed distributions. State the interval method and probability used; a probability label alone does not identify the method. The interval calculation in Abacus summary facades is a separate API contract from the interpretation of a posterior probability.

7. Interpret signs and practical thresholds

A 94% posterior interval above zero is not generally equivalent to rejecting a null hypothesis at a 6% significance level. Posterior probabilities do not provide that frequentist error-rate guarantee. See the conditional interval interpretation.

First inspect the prior support. With a positive continuous prior such as a HalfNormal, $P(\beta > 0 \mid y)=1$ by construction. An interval above zero does not show that the data discovered positivity. Increasing the prior scale cannot introduce negative support. See Bayesian Priors.

Where the model permits the relevant alternatives, report an interval and, if useful, posterior probability relative to a prespecified practical threshold. For example, the probability that a response exceeds a meaningful minimum addresses magnitude within the fitted model. Compare it with the prior probability and assess prior sensitivity. Report probabilities estimated from MCMC draws with appropriate Monte Carlo precision, not as exact calculations.

If an interval includes zero, the estimate is inconclusive as to sign at that interval probability. If it rules out effects of practical importance, state that narrower claim and the threshold. None of these summaries establishes causal identification on its own.

8. Review the evidence before interpretation

  1. Confirm that diagnostics describe retained draws and that the required evidence is available. Missing evidence is not passing evidence.
  2. Investigate retained divergences and flagged R-hat, bulk/tail ESS, energy or tree-depth diagnostics. Inspect trace or rank plots for the same run.
  3. Check Monte Carlo precision for the summaries that will be reported. Increase sampling only where it addresses the diagnosed problem.
  4. Separately assess predictive checks, prior sensitivity, model assumptions and causal identification before using outputs for decisions.

Report the posterior summary, interval method and probability, computational limitations and substantive assumptions. A computational screen can support use of the draws for further analysis; it does not certify the conclusions.

Prior Predictive Checks

If you come from classical econometrics, you are used to checking assumptions after estimation: residual plots, heteroskedasticity tests, outlier influence, and maybe out-of-sample fit. Bayesian workflow adds one earlier question:

Before fitting anything, do my priors imply plausible behaviour for the target variable?

That is what prior predictive checking answers.

1. Why parameter-level priors are not enough

A prior can look sensible when you inspect it in isolation and still imply absurd behaviour once it flows through the whole model.

For example:

  • an intercept prior may look “weakly informative” on paper
  • a channel coefficient prior may look “reasonably positive”
  • a likelihood sigma prior may look “safely diffuse”

But jointly, those choices might imply:

  • weekly revenue that is far above anything you could ever observe
  • negative conversions for a business where the target is always non-negative
  • far more volatility than the real series could possibly have

Classical econometrics rarely forces you to check this explicitly because you usually specify penalties or constraints directly on the coefficient space. Bayesian MMM requires one more layer of discipline: inspect the implied distribution of y, not just the configured priors on the parameters.

2. What prior predictive checking does

Prior predictive checking asks:

If the priors were true, what kinds of target series would this model generate before seeing the actual data?

The workflow is:

  1. Build the model with your chosen priors and structure.
  2. Sample from the prior predictive distribution.
  3. Compare those simulated target draws with the scale and shape of the real target series.

This is not a convergence check and it is not a causal test. It is a plausibility check on the model you are about to fit.

3. How Abacus supports it

Abacus exposes prior predictive sampling directly on PanelMMM:

prior = mmm.sample_prior_predictive(
    X=X,
    y=y,
    samples=100,
    random_seed=42,
)

If you want a quick visual check, Abacus also exposes a retained plot surface:

figure, axes = mmm.plot.prior_predictive(var=mmm.output_var)

In the structured runner, this is Stage 10, the preflight stage. The pipeline writes:

  • 10_pre_diagnostics/prior_predictive.nc
  • 10_pre_diagnostics/prior_predictive.png

Abacus currently gives you the sampled draws and the plot. It does not apply an automatic plausibility score or a hard pass/fail gate for you.

4. What to look for

A useful prior predictive check is not about matching the data exactly. That would defeat the point of a prior. The question is whether the implied target behaviour is at least in the right universe.

Look for the following.

Level

Do the simulated draws live on roughly the same order of magnitude as the observed target?

If your historical weekly revenue is in the low millions, prior predictive draws in the billions are a red flag.

Dispersion

Is the implied volatility remotely plausible?

If the prior predictive distribution is much wider than the observed series, your likelihood sigma or contribution priors are probably too loose.

Sign and support

Does the model imply values that violate business reality?

For example:

  • negative conversions
  • implausibly negative revenue
  • large oscillations around zero for a strictly positive KPI

These are often signs that the prior scale is too permissive relative to the data scaling and likelihood choice.

Time pattern

Do the implied trajectories look structurally plausible?

You are not looking for a perfect seasonal pattern before fitting, but you should ask whether the prior predictive draws look like something that could have come from your business rather than from a random-number generator with no economic interpretation.

5. Common failure modes

Several practical pathologies show up repeatedly.

The intercept is too loose

A very wide intercept prior can dominate the prior predictive distribution, especially when the target has been scaled but the intercept prior is still too diffuse for the transformed space.

The likelihood sigma is too loose

If the prior predictive draws look far too noisy, the problem is often not the media priors at all. It is the observation model allowing implausibly large residual variance.

Media transformation priors are too permissive

Adstock and saturation priors that allow unrealistically persistent carryover or unrealistically steep response can imply contributions that are wildly too large before the data has had any say.

Flexible baseline terms are too unconstrained

Time-varying intercepts, seasonality, events, and other additive effects can all inject structure into the prior predictive distribution. If those priors are too loose, the target series can become implausibly volatile or pattern-heavy before fitting.

6. What to do when the prior predictive check looks bad

Do not proceed directly to posterior interpretation. Fix the model first.

Typical remedies:

  • tighten the intercept prior
  • tighten the likelihood sigma prior
  • make media priors more weakly informative in the economically plausible region rather than completely diffuse
  • reduce unnecessary model flexibility before the data has justified it
  • check whether your scaling choices make the configured priors too wide or too narrow on the model scale

This is the Bayesian analogue of catching a broken specification before you start arguing about p-values.

7. What prior predictive checks do not tell you

Passing a prior predictive check does not mean:

  • the model is causally identified
  • the model will fit well
  • the posteriors will converge cleanly
  • the attribution decomposition will be trustworthy

It only means the configured priors do not imply obviously absurd target behaviour before seeing the data.

You still need:

8. Practical recommendation

Treat prior predictive checking as a standard pre-fit step, not as an optional extra for purists.

In Abacus terms, the workflow should usually be:

  1. Specify the model and priors.
  2. Run sample_prior_predictive(...).
  3. Inspect the implied target behaviour.
  4. Revise the priors if needed.
  5. Fit only once the prior predictive behaviour is broadly plausible.

That sequence is usually cheaper than fitting a badly specified Bayesian MMM and then discovering that the posterior is unstable for reasons you could have caught before sampling.

Posterior Predictive Checks

Posterior predictive checking asks a simple question:

After fitting the model, can it reproduce the main features of the observed data?

For a classically trained econometrician, this is the Bayesian analogue of residual diagnostics, fitted-versus-observed checks, and out-of-sample sanity-checking, but with one important difference: the checks are based on the full posterior distribution, not a single point estimate.

1. What the check actually is

After fitting, you sample from the posterior predictive distribution:

post = mmm.sample_posterior_predictive(
    X=X,
    progressbar=False,
    random_seed=42,
)

Conceptually, each posterior draw says:

  • here is one plausible parameter vector
  • given that parameter vector, here is one plausible target path

If the fitted model is adequate, the observed data should look like a credible member of that posterior predictive family.

2. Why this matters

A model can have:

  • clean MCMC diagnostics
  • seemingly sensible coefficient signs
  • elegant priors

and still fail to reproduce basic features of the target series.

Posterior predictive checks catch that mismatch.

This matters because a model that cannot reproduce the observed target well enough is usually not ready for:

  • decomposition narratives
  • ROI or CPA interpretation
  • budget optimisation
  • strong causal storytelling

3. How Abacus supports it

Abacus exposes posterior predictive sampling directly:

post = mmm.sample_posterior_predictive(
    X=X,
    progressbar=False,
    random_seed=42,
)

It also exposes retained plotting helpers such as:

figure, axes = mmm.plot.posterior_predictive(var=[mmm.output_var])
residual_figure, residual_axes = mmm.plot.residuals_over_time(hdi_prob=[0.94])

In the structured runner, Stage 30 assessment writes a fuller set of artefacts:

  • 30_model_assessment/posterior_predictive.nc
  • 30_model_assessment/posterior_predictive.png
  • 30_model_assessment/posterior_predictive_summary.csv
  • 30_model_assessment/observed.csv
  • 30_model_assessment/fitted.csv
  • 30_model_assessment/fit_timeseries.png
  • 30_model_assessment/fit_scatter.png
  • 30_model_assessment/residuals.csv
  • 30_model_assessment/residuals_timeseries.png
  • 30_model_assessment/residuals_hist.png
  • 30_model_assessment/residuals_vs_fitted.png

That assessment stage is the closest Abacus comes to a retained, systematically-produced posterior predictive diagnostics bundle.

4. What to inspect

Observed versus fitted over time

Start with the time-series overlay.

Ask:

  • Does the fitted mean track the major movements in the target?
  • Are the predictive intervals wide enough to cover the observed series reasonably often?
  • Does the model systematically lag turning points or seasonal peaks?

If the observed line keeps sitting outside the predictive interval in structured ways, the model is missing something systematic rather than merely being noisy.

Residual structure

Residuals should not show strong unresolved patterns.

In practice, look for:

  • long runs of positive residuals followed by long runs of negative residuals
  • clear seasonality left in the residuals
  • residual variance increasing with fitted values
  • one panel slice fitting much worse than the others

The presence of structure in the residuals usually means the model is still under-specified for the data.

Scatter of fitted versus observed

The fitted-versus-observed scatter is not a formal test, but it quickly shows:

  • compression toward the mean
  • systematic underprediction at high values
  • systematic overprediction at low values

This is the Bayesian cousin of the fitted-value plots you would inspect after a classical regression.

5. What “good” posterior predictive behaviour looks like

A good posterior predictive check does not mean the model matches every wiggle exactly.

You are looking for something more practical:

  • the main level and variation are captured
  • the observed series falls inside plausible predictive ranges often enough
  • residuals are not strongly structured
  • panel slices are not failing in obviously asymmetric ways

The question is whether the model is adequate for interpretation, not whether it is perfect.

6. What posterior predictive checks cannot prove

This is the most important warning.

A model can pass posterior predictive checks and still fail as a causal model.

Why? Because posterior predictive checks evaluate prediction of the target, not causal attribution of the components.

Two models can predict sales equally well while assigning very different shares of those sales to:

  • baseline
  • media
  • controls
  • seasonality
  • events

That is why posterior predictive checking must be paired with:

7. Common failure patterns

The model is too rigid

If the fitted line misses broad movements or regime changes, the model may need more structural flexibility, for example in trend, seasonality, controls, or events.

The model is too flexible in the wrong place

You may see good in-sample fit but strange residual behaviour or unstable attribution because the model is fitting noise through components that should remain more constrained.

Media is carrying baseline structure

If media spend is strongly correlated with time patterns, the model may let media soak up baseline variation that should have been handled by intercept, seasonality, controls, or other additive structure.

Baseline is carrying media structure

The reverse can also happen: a very flexible baseline can absorb variation that you would otherwise attribute to media.

8. What to do when checks fail

If posterior predictive checks look bad, resist the temptation to jump straight to interpreting coefficients anyway.

Instead:

  1. Check convergence first.
  2. Inspect residual structure rather than only aggregate fit.
  3. Revisit baseline specification, controls, seasonality, events, and media transformation choices.
  4. Refit and compare again.

In other words, use posterior predictive checking as a model-development tool, not just as a reporting plot.

9. Practical recommendation

In Abacus, the robust sequence is:

  1. Run prior predictive checks before fitting.
  2. Fit the model and verify MCMC diagnostics.
  3. Run posterior predictive checks and inspect residuals.
  4. Only then move to contributions, optimisation, or causal interpretation.

That order mirrors how a careful econometrician would already work, except that the Bayesian workflow makes the predictive-check step much richer and more honest about uncertainty.

Model Comparison

Choose the prediction task before choosing a score. Abacus reports leave-one-out (LOO) and WAIC diagnostics from the stored likelihood, but those scores do not establish causal attribution or automatically assess forecasting.

What ELPD measures

For held-out units indexed by $i$, LOO estimates the expected log predictive density (ELPD) using contributions of the form $\log p(y_i \mid y_{-i})$. ArviZ reports their sum, elpd_loo, not their average. Higher values indicate better predictive performance for the same scoring task and measure. The absolute value depends on the target scale and number of held-out units.

Pareto-smoothed importance sampling (PSIS) approximates these LOO calculations from a fitted posterior, avoiding a refit for every unit when the approximation is reliable. It still requires adequate posterior sampling and suitable importance weights. For the Abacus accessors and report fields, see Diagnostics.

Check comparability before ranking models

Require the same outcome definition and scale, likelihood measure, held-out unit, observations and prediction task. Matching the input CSV is insufficient.

Abacus preset Stored likelihood basis / held-out unit
time_series Outcome level space, with pointwise contributions over dates
fe Within-unit orthonormal contrasts; the outcome-level unit intercept is removed from this likelihood
cre Marginal likelihood for each complete unit block, integrating over the random unit intercept

Use the estimator manifest and the relevant estimator specification to check the actual contract. The compact Bayesian-criteria report distinguishes FE contrast_space from level_space; that field alone does not distinguish CRE unit-block scoring from time-series date scoring.

For example, two time-series specifications fitted to the same dates and outcome scale can be compared if both use the same likelihood measure and acceptable LOO approximations. Comparing FE contrast-space ELPD directly with CRE level-space ELPD is invalid even when both models use the same panel rows. Their predictive densities score different quantities. Changing the target from y to log(y) also requires reconciling the density measure before any comparison; a label change is not sufficient.

Interpret differences and approximation diagnostics

az.compare(...) reports score differences and uncertainty based on the pointwise contributions. Only supply models that satisfy the comparability conditions above. Examine the size and practical relevance of the difference, its estimated standard error and the observations driving it.

A difference small relative to its uncertainty is inconclusive as to which model predicts better. It does not establish equivalence. A two-standard-error rule is not a universal significance test, particularly with few held-out units, dependence or influential observations. Simplicity and interpretability can guide a decision when predictive evidence is inconclusive, but record that as a decision criterion rather than a demonstrated equality of performance.

Check Pareto-k before relying on PSIS-LOO. The diagnostic threshold depends on the number of draws, with 0.7 an upper cap on the usual threshold. Values from 0.7 to 1 indicate unreliable estimates with potentially substantial bias; values at or above 1 indicate a more severe failure. Abacus’s fixed counts above 0.7 and 1 are summaries, not a complete sample-size-dependent reliability assessment. Inspect the ArviZ warnings and pointwise diagnostics too. See the primary PSIS diagnostic reference.

For problematic observations, investigate the data and model, then consider explicit refits or a suitable K-fold/blocked validation design. This guide does not assume an Abacus or ArviZ moment-matching convenience API. Switching to WAIC does not by itself resolve unreliable importance sampling or influential observations.

Assess future prediction with held-out future data

Ordinary LOO can condition on observations later than the omitted date. That is a different task from forecasting without future outcomes. For Abacus’s separate training-prefix fit and later holdout window, use Blocked Holdout Validation. A single terminal block assesses that window; it is not a rolling-origin study or a guarantee for every future horizon. The leave-future-out case study explains the distinction from ordinary LOO.

Keep prediction, model adequacy and attribution separate

Use posterior predictive checks to inspect features relevant to the application, such as volatility and residual time structure. Relative score improvements do not imply that either model is adequate. Conversely, visually similar predictive fits can conceal very different channel contributions.

AIC, BIC, Bayes factors and predictive scores address different objectives and assumptions; they are not interchangeable replacements. In particular, plugging a posterior mean into a likelihood and applying a nominal parameter penalty does not automatically recover the usual AIC/BIC justification.

Assess attribution using the identifying assumptions, prior/baseline sensitivity and relevant external evidence. LOO cannot establish causal identification. See Causal Identification and Baseline vs Media Trade-Offs.

Causal Identification

An Abacus fit estimates quantities conditional on the specified likelihood, priors and data. Interpreting a media contrast as the effect of an intervention requires a separate identification argument. Good predictive fit, stable sampling and narrow posterior intervals do not supply that argument.

Define the intervention and identifying variation

State the channel, spending change, dates, population and outcome for the causal question. Specify how carryover and changes in other channels enter the contrast. Then explain why the observed variation identifies that effect.

For observational MMM, consequential assumptions include:

  • an adequate control set for common causes of media and outcomes, without conditioning on mediators or colliders that invalidate the intended effect;
  • enough relevant variation to learn the response in the spending range of interest, rather than relying entirely on extrapolation;
  • consistent treatment/outcome definitions, and appropriate treatment of interference, carryover and anticipation;
  • an adequate response, baseline and error specification for the estimand.

For example, spending that responds to expected demand can be associated with higher sales even without a media effect. Adding a smooth trend does not necessarily account for that demand information. A control associated with both variables is useful only if its causal role and measurement justify conditioning on it. These assumptions require subject-matter evidence; residual checks alone cannot establish them. See the primary discussion of MMM causal assumptions.

The named FE estimator removes time-invariant unit effects from its slope likelihood. CRE adjusts for its declared transformed between-unit summaries. Neither removes arbitrary time-varying confounding. See Choose an Estimator and Baseline vs Media Trade-Offs.

Assess designs by their assumptions

There is no universal ranking of methods that substitutes for assessing the design and estimand.

Design Key distinction
Randomised experiment Assignment supports an identification argument for the specified treatment contrast; adherence, missing outcomes, interference and the analysis population still matter
Matched-market study without random assignment Matching does not create randomisation; identification depends on the design’s comparability and counterfactual assumptions
Instrumental variables Requires relevance, independence and exclusion; some estimands also require monotonicity or further structural assumptions
Difference-in-differences Requires an appropriate untreated-trend assumption and treatment-timing conditions; compatible pre-trends do not prove parallel counterfactual trends
Regression discontinuity Identification near a cutoff relies on the assignment mechanism and continuity or local-randomisation assumptions; extrapolation requires more

Diagnostics can challenge aspects of a design. They do not generally prove exclusion, absence of confounding or unobserved counterfactual behaviour. Effects identified for different populations or interventions are not interchangeable merely because each estimate is credible for its own task.

Use calibration for the evidence it supplies

Abacus exposes mmm.add_lift_test_measurements(...) and mmm.add_cost_per_target_calibration(...) on a built model. Both are unavailable for the named FE, CRE and release-gated RE presets. Follow Calibration for supported inputs and workflow. EventAdditiveEffect adds an event component; it does not attach experimental lift evidence.

Before calibration, align the external estimate with the model’s channel, units, spending contrast, time window and outcome. Assess the design’s identification, uncertainty and relevance to the modelling population. Avoid counting the same evidence twice without an appropriate dependence model.

Calibration conditions the fitted response on the supplied evidence under the calibration model. Conflicting observations and calibration data need investigation; their combination is not guaranteed to remove bias. A test for one channel and window does not identify every other channel or establish transportability to a different spending range. Record these limits alongside the calibrated result.

Interpret scenarios and optimisation conditionally

A scenario evaluates a specified spending plan under the fitted response functions and planning assumptions. Optimisation selects a plan for a chosen objective and constraints. Neither operation independently verifies that the intervention will produce the predicted change.

A channel ranking at historical spend is insufficient for allocation. The optimum depends on marginal responses across feasible spending levels, saturation, carryover and constraints. Two models can rank historical returns in the same order yet recommend different allocations. Proportional bias in some reported channel estimates does not establish that their full marginal response functions are correct.

Report the estimand, units, assumptions, supported spending range and uncertainty. Posterior intervals condition on the fitted model; they do not automatically include uncertainty about omitted confounding, model choice or future structural changes. Use sensitivity analyses and relevant experimental evidence to assess those risks. See Interpreting Optimisation.

Baseline vs Media Trade-offs

One of the most confusing experiences in MMM is this:

  • two specifications can fit the target series almost equally well
  • both can have acceptable diagnostics
  • yet they can assign very different amounts of the target to media versus baseline

This is not necessarily a bug in the software. It is a structural feature of the problem.

This page explains how that trade-off appears in Abacus and why you should expect it.

1. The decomposition problem

At a high level, Abacus builds the expected target from several additive components.

In the retained PanelMMM build path, the mean function can include:

  • intercept_contribution
  • channel_contribution
  • control_contribution, if you configure control_columns
  • mundlak_contribution, if use_mundlak_cre=True
  • yearly_seasonality_contribution, if yearly_seasonality is enabled
  • additional additive effects you attach before build, such as events or trend effects

The likelihood sees the sum of these pieces, not a directly observed “ground-truth baseline” and “ground-truth media” split.

That means the total fit can be easier to identify than the decomposition.

2. Why the trade-off exists

Suppose revenue rises every December and TV spend also rises every December.

Several stories can fit the same sales data reasonably well:

  • December uplift is mostly seasonality
  • December uplift is mostly TV
  • December uplift is partly both

If the model includes both a seasonal term and media terms, they will compete to explain the same observed movement.

This is the core baseline-versus-media trade-off:

the data often identify total explained variation better than they identify which component deserves the credit

Classical econometricians already know this as collinearity and omitted-variable competition. Bayesian MMM does not make that problem disappear. It makes the uncertainty around it more explicit.

3. What counts as “baseline” in Abacus

In Abacus, the baseline side comes from the terms you specify inside the PyMC graph.

Depending on configuration, that can include:

  • a static intercept
  • a time-varying intercept
  • yearly Fourier seasonality
  • controls
  • events
  • trend-like additive effects
  • Mundlak CRE adjustments in panel settings

So when people say “baseline absorbed the effect”, they usually mean one or more of those components, not a separate external decomposition engine.

4. How media can lose attribution

Media can lose attribution when the non-media side of the model is too good at explaining the same movements.

Common cases:

  • a flexible time-varying intercept captures medium-run swings that media could also explain
  • strong seasonal terms absorb repeating peaks that coincide with campaign timing
  • control variables proxy for media timing or market conditions too strongly
  • event effects explain demand spikes that were previously being picked up by channel coefficients

In each case, the model may still predict well. The question is how the variation is partitioned.

5. How media can steal attribution from baseline

The reverse failure is also common.

If the baseline side is under-specified, media channels can absorb variation that is not truly incremental media response.

Examples:

  • missing seasonality leaves recurring annual structure for media to explain
  • missing controls leave competitor, pricing, or macro effects for media to explain
  • missing events leave spikes for channels to absorb
  • insufficient baseline flexibility forces media to act as a trend proxy

This usually inflates media contribution and makes optimisation outputs look better than they should.

6. Why good fit does not settle the argument

You might hope that whichever specification predicts better must also have the more trustworthy attribution split.

Unfortunately, that does not follow.

A model can reproduce the observed target series very well while still having ambiguous attribution. Predictive adequacy is necessary, but it is not enough to identify the correct media decomposition.

That is why:

7. Signs that the trade-off is driving your result

Be cautious when you see any of the following:

  • very similar model fit with materially different channel contributions
  • large channel swings after adding or removing a seasonal or trend term
  • media ROI rankings that flip after adding controls or events
  • one highly flexible baseline term dominating decomposition while media contributions collapse
  • implausibly smooth media contributions paired with a very wiggly baseline, or vice versa

These are not proofs of misspecification, but they are strong prompts for sensitivity analysis.

8. What to do in practice

A disciplined Abacus workflow is usually better than trying to argue theoretically about the “right” split in the abstract.

Recommended approach:

  1. Start with a specification that has the minimum baseline structure you can defend.
  2. Add seasonal, control, event, or time-varying terms only when you can justify them substantively or diagnostically.
  3. Refit and compare decomposition stability, not just target fit.
  4. Report instability when attribution changes materially across defensible specifications.
  5. Where possible, bring in external evidence such as lift tests or calibration.

The important point is not to force one narrative prematurely. It is to show which attribution conclusions remain stable after reasonable specification changes.

9. Abacus-specific interpretation

In Abacus, you should treat the decomposition outputs as conditional on the configured structure:

  • the chosen controls
  • whether yearly_seasonality is on
  • whether the intercept is time-varying
  • whether media effects are time-varying
  • whether you added events or other additive effects
  • whether use_mundlak_cre=True

Change the structure, and the attribution can change even when predictive fit does not move much.

That is normal. It is the software telling you where the data alone are not decisive.

10. Bottom line

Baseline-versus-media trade-offs are unavoidable in MMM because the observed target only reveals the sum of the contributing processes.

Abacus makes this explicit by fitting all configured terms inside one additive Bayesian graph. That is a strength, but it also means you need to read the decomposition as a conditional statement:

given this model structure, priors, and data, this is the most plausible attribution split

That is much more defensible than pretending the split is uniquely observed in the data.

Mundlak Specification Test

Background

Classical panel econometrics uses a Mundlak specification test to assess the mean-independence restriction behind a random-effects (RE) model. In its usual form, the test evaluates whether the coefficients on declared unit-level regressor summaries are jointly zero.

A rejection is evidence against that particular RE restriction under the test assumptions. A failure to reject is not proof that RE is adequate: weak variation, collinearity, finite samples, or an incomplete summary basis can leave the test uninformative.

Why Abacus Does Not Reproduce the Frequentist Test

Abacus fits Bayesian models. It does not attach an asymptotic Wald test or a chi-squared reference distribution to the Mundlak coefficients. Posterior inference on those coefficients answers a different question and depends on the declared priors and summary basis.

Two interpretations must be avoided:

  • A posterior interval containing zero does not establish that the RE mean-independence assumption is adequate.
  • A posterior interval excluding zero indicates a conditional association with the declared summaries. It does not identify the amount of confounding or establish causal identification.

This distinction is especially important in marketing mix modelling, where media transformations are estimated and the available between-unit variation may be weak.

Legacy Mundlak Surface and Named CRE Preset

use_mundlak_cre=True is the retained low-level panel surface. It adds legacy Mundlak terms to an unlabelled dimensioned PanelMMM. It is not an alias for the named cre estimator preset.

The named CRE preset has a separate, explicit contract. Its media summaries use the declared transformed exposure basis and it records estimability evidence. The released v1 surface retains explicit limits on prediction and unsupported downstream operations.

Do not transfer an interpretation or diagnostic result from one surface to the other without checking the actual fitted summary basis.

What to Inspect

Posterior summaries

For a legacy Mundlak fit, inspect the coefficients and their joint posterior geometry:

import arviz as az

az.summary(
    mmm.idata,
    var_names=["gamma_channel_mundlak", "gamma_control_mundlak"],
)

Treat these summaries as evidence about associations conditional on the fitted model. Check effective sample size, R-hat, posterior correlations, prior sensitivity, and the amount of within- and between-unit variation before interpreting their magnitude.

Estimability and sensitivity

Before relying on the adjustment:

  1. Confirm that the declared summaries have non-zero between-unit variation.
  2. Inspect rank, collinearity and condition-number diagnostics.
  3. Compare posterior results under defensible prior alternatives.
  4. Check that substantive conclusions are not driven by one summary-basis choice.
  5. Use prior and posterior predictive checks to detect implausible model behaviour.

Predictive comparison

A predeclared held-out comparison can test whether adding the adjustment improves prediction for the intended forecasting task. Use a split that respects panel and temporal dependence. Predictive improvement does not by itself validate the identifying assumption or convert observational associations into causal effects.

Summary

Evidence Supported conclusion Unsupported conclusion
Adjustment interval includes zero The data and prior do not clearly separate the coefficient from zero RE is adequate
Adjustment interval excludes zero Association with the declared unit-summary basis Identified confounding or causal correction
Predictive comparison improves Better prediction for the declared holdout task Correct causal structure
Estimability diagnostics pass No detected defect under the implemented screens Global or posterior-wide identification

Abacus should therefore retain explicit posterior, estimability, sensitivity and predictive evidence. It should not turn the Bayesian adjustment into a binary RE-versus-CRE adequacy test.

References

  • Mundlak, Y. (1978). “On the Pooling of Time Series and Cross Section Data.” Econometrica, 46(1), 69–85.
  • Vehtari, A., Gelman, A., & Gabry, J. (2017). “Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC.” Statistics and Computing, 27(5), 1413–1432.