This section collects longer-form answers to recurring MMM, Bayesian, and
panel-econometrics questions that come up when practitioners move from
classical econometrics to PanelMMM.
The pages are written for technical readers who already understand regression,
panel data, and causal inference, but want the Abacus framing. Start with the
practical questions to connect those concepts to an analyst’s workflow.
Is my data suitable for Abacus, and what must I prepare?
Check two things before fitting: whether the data satisfy Abacus’s input
contract, and whether they contain enough relevant variation to answer your
question. A correctly formatted dataset can still support weak or ambiguous
channel estimates.
How much history do I need?
There is no universal minimum number of weeks, markets or observations that
makes an MMM reliable. Assess the history against the proposed model:
Does it cover the seasonal patterns and business regimes you want to model?
Do channels change independently enough to separate their responses?
Is there information about different exposure levels, including the range
relevant to the proposed decision?
Is there enough earlier history for the chosen adstock horizon and enough
data left after reserving a meaningful holdout?
For a panel estimator, is there usable variation within units and, where
required, between units?
Three years of channels that always move together may contain less useful
channel-specific information than a shorter period with distinct changes.
Adding markets does not automatically add independent information either.
Priors can regularise weakly informed parameters; they do not create the
missing variation. See Bayesian Priors.
Should my media inputs be spend, impressions or something else?
Use a consistently measured exposure variable that matches the response you
want to model. Abacus applies adstock and saturation to the channel columns
you supply. It does not obtain a separate monetary spend series or convert
impressions into money for you.
The choice also determines which downstream interpretation is valid:
Channel input
Outcome
Contribution divided by input means
Monetary spend
Revenue in the same currency
Model-conditional ROAS
Monetary spend
Conversion count
Conversions per unit of spend; its reciprocal is cost per conversion
Impressions
Revenue
Revenue per impression, not ROAS
The standard budget workflow treats allocated channel-input quantities as
spend. Do not pass a monetary budget into a model trained on impressions and
assume Abacus supplies the price conversion. That requires an explicit,
reviewed mapping and a compatible planning workflow.
Keep currency, tax treatment, gross/net definitions and attribution dates
consistent. A change in measurement can resemble a change in response.
What table must I provide?
For Python fitting, supply X as a DataFrame and y as a row-aligned Series or
one-dimensional array. Keep the target out of X; include the date, channel,
control and declared panel-dimension columns. A combined CSV is supported by
the YAML and runner workflows, which separate the target for you.
For an aggregate time series, use one observation per date. For a panel,
provide exactly one row for each expected date and panel-coordinate
combination. Multiple panel dimensions require the declared rectangular
grid; they do not represent arbitrary nested or incomplete groups.
Use a consistent time frequency and align the outcome, media and controls to
the same periods. l_max=8 refers to eight input periods, including the current
period, not necessarily eight weeks. Document aggregation rules: sums can
make sense for spend and revenue, while a price or rate needs a justified
aggregation method.
No. Abacus requires observed values for the channel, control and target cells
you supply; it does not silently replace missing measurements with zero.
Absent panel rows and duplicate rows also need resolution before fitting.
For example, consider two markets in one week:
Market
Recorded TV spend
Interpretation and action
North
0
Keep zero if the source confirms no TV spend occurred
South
Missing
Investigate the missing measurement; do not infer zero spend
If South’s sales are also missing, adding a row with zero sales invents an
observation. Recover the measurement or choose and document a defensible
treatment of the missing data. Any imputation introduces assumptions whose
effect on the analysis should be assessed. If you change the analysis window
or units, check the panel contract again.
What does Abacus preprocess automatically?
Abacus scales the target and channel inputs before constructing the model.
It does not automatically scale controls, choose a control set, reconcile
currencies, adjust for inflation or perform domain-aware missing-data repair.
Default target and channel scaling pools over the panel dimensions. It is not
automatically separate scaling for each market. Check the configured scaling
dimensions before choosing prior magnitudes or comparing parameters.
Supply the media inputs intended for the configured transforms. Do not
pre-adstock or pre-saturate them and then apply the same transformations again
inside Abacus. Check the outcome implications of your complete specification
with prior predictive draws on the appropriate scale.
Should I combine correlated channels or add more controls?
Neither action is an automatic remedy. If two channels always run together,
separate contributions may depend heavily on priors and specification choices.
Consider a combined channel when it represents a coherent exposure and the
decision can be made at that combined level. Combining channels changes the
estimand: it cannot then justify separate channel allocations. Changing media
mix or unit costs can also make the combined response unstable.
If separate estimates are essential, seek additional identifying variation
or relevant experimental evidence, and report the limits of the current data.
Inspect transformed predictors as well as raw correlations; the model learns
from adstocked and saturated exposures.
Choose controls from their role in the outcome and media-assignment process,
not solely because they improve fit. A variable affected by media can block
part of the effect you intended to estimate. More controls can also introduce
collinearity. See Causal Identification.
What should I prepare before the first fit?
Record the target, units, time frequency, panel definition and decision
question. Retain an audited input table, its source and all preprocessing
decisions. Review variation, plausible controls, model scale and priors.
Choose the estimator and reserve the validation window before comparing
preferred results.
Use Blocked Holdout Validation
for a training-prefix refit and later holdout. The supplied holdout media and
controls make this a conditional prediction exercise; it does not establish
that those inputs would have been known in a live forecast. Keep information
from the holdout out of fitted preprocessing and model selection wherever
that information would be unavailable at prediction time.
Which Abacus model should I choose, and what does that choice permit?
Choose the model from your question, data structure and defensible assumptions
before examining preferred channel results. In Abacus, the choice also
determines which prediction, calibration and planning operations are available.
This page describes the accompanying Abacus 3.1.1 library. The
estimator support matrix
is the detailed reference for named presets.
Which model is the starting point for my data?
Situation
Candidate
What you must justify
One aggregate observation per date
Named time_series preset
Temporal variation, baseline, controls and media response specification
A rectangular panel with a deliberately configured parameter and prior structure
Ordinary dimensioned PanelMMM
Which parameters vary by slice, which are shared and whether any hierarchical pooling is specified
A balanced panel where shared slopes should use changes within units
Named fe preset
Sufficient transformed within-unit variation and the remaining time-varying confounding assumptions
A balanced panel requiring an explicit adjustment for persistent unit differences associated with predictors
Named cre preset
The declared transformed between-unit summary basis, its estimability and sufficient within-unit variation
The named FE and CRE presets take one unit dimension, such as geo, containing
multiple units. This is not a claim that they fit only one market. Their
released media slopes and adstock/saturation parameters are shared across
units. Do not interpret a separate regional contribution as a separately
estimated regional response function.
Does setting dims select FE or create partial pooling?
No. dims=("geo",) declares an array axis. On the ordinary PanelMMM surface,
the parameter dimensions and priors determine which quantities are shared or
vary by market. Independent market-indexed priors are not hierarchical
partial pooling; shared parameters are not the named FE likelihood.
Specify a named estimator explicitly when you want its contract. The legacy
use_mundlak_cre=True option on an ordinary panel model is also not an alias
for the named CRE preset. See Panel Dimensions
and the Mundlak FAQ.
How do FE and CRE differ in the information they use?
FE removes persistent unit intercepts from its shared-slope likelihood by
using within-unit contrasts. It cannot learn a slope from a predictor that
does not vary within any unit. A larger market’s consistently higher spend
and sales do not supply the same information as changes within that market.
CRE models unit intercepts together with declared centred between-unit
summaries. In the named Abacus preset, media summaries use fitted transformed
exposures, not simply raw-spend averages. The adjustment is limited to that
declared basis and needs usable between-unit information as well as within-unit
variation.
For example, suppose larger markets always spend more and sell more. FE asks
what the within-market changes reveal under its model. CRE additionally
represents the specified relationship between persistent market differences
and predictor summaries. If spend barely changes within markets, neither
preset creates the missing within-market information.
Neither removes arbitrary time-varying confounding. The released named
presets also do not support common categorical time effects. Read the
FE specification and
CRE specification
before importing assumptions from another panel package.
Why is the declared re preset unavailable?
Implementing and qualifying the named random-effects preset was a lower
priority because its assumptions are often difficult to defend in marketing
applications.
Standard RE requires the unobserved persistent unit effect to have conditional
mean zero given the included predictor history. Marketing budgets commonly
reflect market size and expected baseline demand, which also affect sales.
When the model does not adequately account for these factors, the RE
restriction is implausible.
Correlation between spend and observed market size is not itself a
violation if the specification adequately accounts for market size. The
restriction concerns the remaining unobserved unit effect; it is not a rule
that raw spend and market size must be uncorrelated.
Abacus recognises estimator.type: re in configuration, but the named preset
has not passed the required implementation checkpoint for public release.
Building it raises EstimatorReleaseGateError before fitting. Internal
implementation work does not make it a released estimator. This is an
intentional restriction, not an installation problem or a general statistical
objection to random-effects modelling.
Choose another estimator only when its assumptions fit the question. CRE is
not an automatic substitute for every RE analysis.
Can every model predict, calibrate and optimise?
No. For the named presets in this release:
Operation
time_series
fe
cre
Prediction on new dates with supplied inputs
Supported
All fitted units required
All fitted units and frozen fitted CRE summaries required
Historical and manual allocation scenarios
Supported
Supported for all fitted units
Supported with the same fitted-unit and frozen-summary restrictions
Lift-test or cost-per-target calibration
Supported through the ordinary model surface
Unavailable
Unavailable
Fixed-budget optimisation
Supported
Unavailable
Unavailable
Supported operations still require valid inputs and a suitable fitted model.
For ordinary dimensioned PanelMMM, check the configured model and operation
contracts rather than assuming the named-preset table describes every custom
configuration. New dates do not imply support for previously unseen markets.
For panel manual allocations, use a labelled xarray.DataArray or
DataArraySpec covering the required fitted coordinates. A channel-only
dictionary is not a panel allocation. See
Scenario Specifications.
An unsupported-operation error is not resolved by removing the estimator label
from a fitted model. That would discard the contract without making the
operation statistically valid.
Should I use Python, a YAML builder or the pipeline runner?
These are workflow choices rather than competing estimators. Python gives
direct control over construction, fitting and analysis. The YAML builder
constructs a model from configuration for subsequent use. The runner executes
a staged workflow and writes artefacts to disk.
Their configuration surfaces differ. Start with the appropriate
Python,
YAML builder or
runner guide. A successful build or
run does not establish that the chosen model is suitable for the data.
Can I choose the model with the best LOO score?
Only compare scores that describe the same prediction task, observations,
outcome scale and likelihood measure. Abacus’s named FE likelihood scores
within-unit contrasts; CRE scores complete unit blocks with its unit intercept
integrated out. Those scores cannot be ranked against each other as though
they evaluated identical outcome-level observations.
Use the Model Comparison FAQ to
check comparability. Assess computation, predictive adequacy, sensitivity and
identification separately. Choose the model because its contract answers the
question, then report what the evidence permits.
What do contribution, ROAS and scenario results actually mean?
Read each output as a quantity defined by the fitted model, its units and its
evaluation window. A contribution table, revenue prediction and allocation
comparison answer different questions. None independently establishes a
causal media effect.
Is media contribution the same as predicted revenue?
No. In the ordinary additive model, media contribution is the media component
of the fitted mean. The full mean also includes the configured intercept,
controls, seasonality and other additive terms. Posterior predictive outcome
draws additionally include observation uncertainty through the likelihood.
Output
What it describes
What it does not automatically describe
Historical channel contribution
A component evaluated for the observed media path under the fitted model
A measured causal increment
Expected media contribution under a scenario
The model’s media response to the specified input path
Total future revenue or realised sales
Posterior predictive outcome
An outcome draw conditional on supplied inputs and the fitted model
Uncertainty about every possible future input or structural change
Optimised allocation
A solution for the configured objective, budget and constraints
An approved business recommendation
Named FE and CRE have their own likelihood and prediction contracts; do not
transfer every ordinary-model interpretation without checking those contracts.
See Model Overview and
Contributions and Decomposition.
When is a reported ratio really ROAS or CPA?
The efficiency accessors use fitted contributions and the supplied channel
inputs. They do not look up an independent spend series.
Revenue contribution divided by monetary spend in the same currency gives
model-conditional return on advertising spend (ROAS).
Monetary spend divided by conversion contribution gives model-conditional
cost per acquisition/conversion (CPA), with the conversion definition stated.
Revenue contribution divided by impressions is revenue per impression,
even if an output column is labelled ROAS.
For example, £200 of modelled revenue contribution divided by £100 of spend
gives ROAS 2. £200 divided by 1,000 impressions gives £0.20 per impression.
Neither calculation is profit: margins and other costs are separate.
Use original-scale contributions and consistent periods and panel aggregation.
target_type selects an efficiency accessor and label; it does not validate
currencies or convert exposure units. The element-wise accessors return NaN
for a zero denominator. Do not replace that undefined ratio with a favourable
return. See ROAS and Metrics.
How should I aggregate returns and uncertainty?
Define the aggregate quantity first. For a total-window return, sum
contributions within each posterior draw over the intended dates and regions,
then divide by the corresponding total spend. Summarise the resulting draw
distribution. An unweighted average of weekly or regional ratios generally
answers a different question.
For example, spend of £100 and £900 with contributions of £300 and £900 gives
individual ROAS values of 3 and 1. Their unweighted mean is 2, but the combined
return is £1,200 / £1,000 = 1.2.
Similarly, sum contributions within each draw before calculating an interval
for the total. Adding component interval endpoints does not generally give
the total’s credible interval because the components are dependent. Inspect
the aggregation behaviour of the specific accessor you use; a frequency
argument alone does not define the desired ratio estimand. See
Summary and Export.
Why might optimisation favour a channel with lower historical ROAS?
Historical average ROAS describes return over the evaluated exposure path.
Allocation decisions depend on the change in the objective from feasible
changes in spend. With saturation, a channel can have high historical average
return but low marginal response at its current spending level.
For the low-level PanelBudgetOptimizerWrapper, the default objective is
average posterior total_media_contribution_original_scale. It is not
automatically profit, a lower credible bound or a risk-adjusted business
utility. The feasible solution also depends on bounds, time allocation and
carryover assumptions.
A channel at its upper bound may indicate that the objective would prefer
more spend if allowed. It does not establish that the boundary is a validated
commercial optimum. Inspect the solver result, constraint satisfaction,
supported spending range and sensitivity before interpreting it. See
Budget Optimisation.
Total across allocation cells for one model period
ManualAllocationScenarioSpec.allocation
Total over the requested spending window for each allocation cell
FixedBudgetOptimizedScenarioSpec.total_budget
Total over the requested spending window
Preferred runner optimization.budget block
Total-window budget, resolved according to its configured mode
Legacy runner optimization.total_budget
Per-period budget
For eight weekly spending periods, a flat £80,000 window budget corresponds to
£10,000 per period across all allocation cells. Passing budget=80_000 to the
low-level wrapper instead requests £640,000 over those eight periods.
Why do historical and simulated scenarios give different results?
A historical reference uses observed history. A simulated allocation defines
a spending path, which may differ in timing even when its channel totals
match that history. Adstock and saturation make timing consequential.
Check the requested spending window, evaluated response window, incoming lag
history, time distribution and carryover tail. With include_carryover=True,
the response window extends beyond the spending window. A default simulated
scenario uses include_last_observations=False, so historical carryover is
not automatically included.
noise_level controls simulated spend-path variation; it is not the outcome
likelihood’s residual uncertainty. Set it to zero when you need a deterministic
spend path. Read the resulting metadata rather than inferring the scenario
definition from its name.
CurrentScenarioSpec requires overlap with observed dates. It is not a
future-only no-change forecast. See Overview and Workflow.
How do I compare two plans and report uncertainty in their difference?
Define a common response quantity and align the plans’ response horizon,
initial history, carryover treatment and posterior draws. For a fixed-budget
reallocation question, keep total spend the same. An expansion plus
reallocation answers a different question and should be labelled accordingly.
The following is illustrative arithmetic, not a fitted Abacus result:
Plan
Eight-week spend
Mean media contribution over the same response window
Reference allocation replay
£80,000
£160,000
Alternative allocation replay
£80,000
£176,000
The mean difference is £16,000. That table alone supplies no credible interval
for the difference. Compute the alternative minus reference contribution for
each matched posterior draw, then summarise those differences. Do not subtract
the endpoints of the two marginal intervals or infer the difference’s
uncertainty from whether those intervals overlap.
If the alternative instead spends £88,000, report an expansion-plus-reallocation
contrast. If the reference is historical attribution with different incoming
history, re-evaluate a matching reference path before making the controlled
allocation comparison.
ScenarioPlanner.compare(...) concatenates individual scenario summaries;
it does not by itself provide a paired difference interval. The pipeline’s
matched allocation comparison retains separate paired-difference artefacts.
Check the actual output contract in Comparison Outputs
and Output Directory Schema.
What is required before recommending a budget change?
Review computation, prior plausibility, predictive checks and relevant
holdout evidence. Then assess identification, parameter and specification
sensitivity, extrapolation, commercial constraints and the uncertainty in
the decision quantity. A solver success flag or narrow posterior interval
does not replace these checks.
State the response quantity, units, spending and response windows, assumptions
and permitted interpretation with the recommendation. Posterior uncertainty
conditions on the fitted model; it does not automatically include omitted
confounding, model-choice uncertainty or future structural change. See
Causal Identification and
Interpreting Optimisation.
Bayesian Priors
Priors constrain and regularise the model. Their influence depends on their
support and scale, the likelihood and the information in the data. A proper
posterior or stable fit does not by itself establish identification.
Support restrictions and prior scale
Separate two choices:
Support: which parameter values the model permits.
Concentration: how it distributes probability over those values.
A HalfNormal prior gives zero probability to negative values. A LogNormal
prior restricts its parameter to strictly positive values. Increasing either
prior’s scale does not allow the posterior to become negative: the likelihood
cannot create posterior support where the prior assigns none.
Consequently, P(beta > 0 | data) is one by construction for a coefficient
with a positive continuous prior. It is not evidence that the data established
a positive effect. An interval above zero must be interpreted in light of
that restriction. A posterior can still concentrate near zero; whether it
rules out a practically negligible effect is a separate question.
A concentrated prior can reduce variance while introducing bias when its
assumptions are wrong. A diffuse prior can permit implausible values or weakly
identified combinations. Neither choice guarantees accurate estimates, good
sampling or causal validity. “Weakly informative” only has meaning relative
to the parameterisation and input/output scales.
Relation to penalised estimation
Under a specified likelihood, a Normal prior yields a quadratic penalty in
the negative log posterior, and a Laplace prior yields an absolute-value
penalty. This connects posterior modes to ridge and lasso estimates in the
corresponding models. It does not make their full uncertainty calculations
identical.
An unconstrained frequentist estimate does not require a Bayesian prior. A
flat density on the whole real line is improper, and flatness changes under
non-linear reparameterisation. Avoid treating it as an assumption-free
probability distribution.
Specify priors in Abacus
Use Prior from pymc_extras.prior. For example, these transform objects
restrict the response amplitude to be non-negative and the decay to (0, 1):
This is a configuration fragment: pass the objects as adstock and
saturation when constructing PanelMMM. The numerical scales above are
illustrative, not a recommendation for every dataset. Abacus applies media
transforms to scaled inputs, so assess their implications on the outcome
scale too. For model-level overrides and valid YAML syntax, use
Priors and Configuration.
Choose and assess a prior specification
State the support restrictions and their substantive justification.
Check plausible parameter magnitudes in the actual model scales.
Assess the available identifying variation, including correlated media,
persistent unit differences and possible time-varying confounding.
Fit defensible alternative prior specifications and compare the quantities
used for decisions, including contributions and scenario contrasts.
There is no universal threshold in weeks or number of channels that makes
media effects identified or a prior negligible. More observations need not
supply independent variation. External calibration can inform a particular
response, but its relevance depends on the experimental design, estimand and
transport assumptions; see Calibration.
Compare like quantities
Compare each parameter prior with its parameter posterior to inspect
updating. Similar distributions do not uniquely diagnose an excessively
concentrated prior: the likelihood may be weak, compatible with the prior, or
informative about combinations rather than individual parameters. A shifted
posterior also does not show that the prior has ceased to matter.
Compare prior predictive and posterior predictive distributions with
observed outcomes to assess their implications for data. These distributions
include different sources of uncertainty from a parameter distribution and
cannot be substituted for it. Sensitivity analysis is needed to assess how
conclusions depend on the prior, even when the posterior looks concentrated.
Adstock represents carryover; saturation represents a non-linear response to
media exposure. In PanelMMM, the selected transform families and their
priors define the response model. Joint estimation propagates uncertainty in
those parameters, conditional on the specified model; it does not remove the
need to assess identification and functional-form sensitivity.
Geometric adstock and its finite horizon
For the usual lagged convolution, with L = l_max, geometric adstock uses
weights proportional to $\alpha^\ell$ for $\ell=0,\ldots,L-1$. With the
GeometricAdstock object’s default normalize=True, the transformed series is
With normalize=False, omit the denominator. This is a finite convolution,
not an untruncated recursive Koyck model. l_max=4 retains the current period
and three preceding periods. The available history and convolution mode also
matter, especially at the start of a series or prediction window.
The default decay prior is Beta(alpha=1, beta=3). A decay near zero gives
little weight to older exposure; a decay near one retains substantial weight
across the specified horizon. The prior does not ensure that omitted lags are
negligible. Choose l_max using plausible carryover and assess sensitivity to
that choice. The lag count refers to the input frequency, not necessarily
weeks.
Alternative transform families encode different lag-weight shapes. Inspect
their parameterisation rather than assuming that every alternative permits a
delayed peak. See Adstock and Saturation
for configuration and supported classes.
For positive beta and lam on non-negative inputs, the response is concave
and approaches beta. It reaches half that limit at
$x=\log(3)/\lambda$. Increasing lam moves half-saturation towards zero; it
does not create a freely located positive-spend inflection point. This is not
a general sigmoid with a learnable threshold before diminishing returns begin.
The default priors are Gamma(alpha=3, beta=1) for lam and
HalfNormal(sigma=2) for beta. Their implications depend on the model’s
scaling: the transform receives scaled inputs, and beta sets the response
amplitude on the model scale. See
Scaling and Preprocessing
and Bayesian Priors. Positive support
is a modelling restriction, not evidence learned about the sign.
Joint estimation and omitted uncertainty
Abacus estimates the uncertain transform parameters jointly with the other
model parameters. In this logistic specification, beta is the response
amplitude; it is not an additional coefficient to be multiplied by a separate
generic media slope.
Fixing transform parameters before fitting conditions on those choices.
Joint estimation can propagate their posterior uncertainty, but still
conditions on the selected transform families, lag horizon, controls and
baseline. Neither approach guarantees narrower, wider or better-calibrated
intervals in every dataset. Assess sensitivity in the quantities used for
decisions, not only in parameter summaries.
Transformation order
adstock_first=True, the PanelMMM default, applies saturation to accumulated
media exposure. adstock_first=False applies saturation within each period
before carrying the response across periods. These operations generally do
not commute.
Choose the order from the response mechanism and assess its implications;
channel names alone do not determine it. Normalisation and lag length affect
the result as well. A good fit under either order does not prove that its
media decomposition is correct. See
Baseline vs Media Trade-Offs.
HSGP
A Hilbert space Gaussian process (HSGP) approximates a Gaussian process with
a finite set of basis functions. In Abacus it provides a regularised function
for components such as a smooth baseline. Basis size, covariance assumptions
and hyperpriors all affect what that component can represent.
Basis size and model flexibility
The m setting controls the number of retained basis functions. Conditional
on the GP hyperparameters, Abacus gives the basis coefficients Normal priors
whose scales depend on the covariance’s spectral density. These priors can
shrink high-frequency terms; they do not generally set them exactly to zero.
The number of basis coefficients is not an OLS residual-degrees-of-freedom
calculation. Regularisation can make effective flexibility smaller than the
basis count, but the data do not automatically select a uniquely correct
amount of flexibility. Poorly constrained hyperparameters and baseline/media
trade-offs can remain even when fitting succeeds.
Select and check the approximation
HSGP.parameterize_from_data(...) recommends m and the boundary extent
from the supplied data and lengthscale settings. It provides a starting point,
not proof that the approximation is adequate for the fitted posterior.
A basis that is too small can miss variation permitted by the covariance.
Increasing m can change the fitted function until the approximation is
adequate, and increases computation. There is no guarantee that m=50 and
m=500 produce identical curves. Assess both approximation settings and
hyperprior sensitivity; do not choose them solely to obtain a preferred media
result. Riutort-Mayol et al. describe basis
and boundary selection and diagnostics for approximation adequacy.
Baseline and media remain competing explanations
A regularised baseline can still explain variation that also aligns with
media. Orthogonality among mathematical basis functions does not imply
orthogonality to the observed media design. Shrinkage changes the allocation
of variation; it does not remove omitted confounding or identify the causal
media contribution.
Check whether substantive results change across plausible baseline and prior
specifications, as well as whether the total fit changes. See
Baseline vs Media Trade-Offs.
Choose between Fourier and HSGP seasonality
Component
Assumption and practical consideration
YearlyFourier
A finite seasonal basis with the configured coefficient priors; order controls retained harmonics
HSGP
A finite approximation to a non-periodic GP; covariance and hyperpriors control smoothness and scale
HSGPPeriodic
A finite periodic GP approximation; the seasonal function repeats at its configured period
The retained HSGPPeriodic does not introduce slowly drifting seasonal
coefficients. Modelling changing seasonal shape requires an explicit model
for that change; choosing this class alone does not provide one.
HSGP uses a basis representation rather than the full observation covariance
factorisation of an exact GP. Its cost still depends on basis size, model
structure and sampling behaviour. Neither HSGP nor Fourier is universally
superior. Use Seasonality and Trends
for supported configuration, then assess predictive behaviour and attribution
sensitivity for the intended task.
Model dated events explicitly
A holiday indicator and a smooth dated-event basis encode different temporal
shapes. Choose according to the event’s expected duration and the data; a
smooth build-up and decay is not always more realistic than an indicator.
For event attachment, use the example and prerequisites in
Additive Effects and Events.
Attach the effect to an unbuilt model, then build and fit with matching X
and y. Event effects add explanatory components; they are not lift-test
calibration. Check that reference’s persistence limitation before saving a
model containing events.
MCMC Diagnostics
Use MCMC diagnostics to assess whether retained simulation draws support the
posterior summaries you want to report. They concern Monte Carlo exploration
and precision. They do not establish that the model is correctly specified,
that media effects are identified, or that predictions will generalise.
For Abacus methods, report fields, threshold comparisons and unavailable
states, use the canonical diagnostics guide.
1. What the sampler does
Markov chain Monte Carlo (MCMC) approximates posterior expectations and
probabilities using dependent simulation draws. Abacus’s usual NUTS workflow
uses Hamiltonian Monte Carlo trajectories with an adaptive trajectory length.
Warmup adapts sampling settings; retained draws are used for posterior
summaries. A larger retained draw count does not by itself establish adequate
exploration.
For example, two chains with 2,000 retained draws each provide 4,000 parameter
vectors. Means, intervals and estimated posterior probabilities derived from
those vectors have Monte Carlo error. Keep that simulation error distinct from
posterior uncertainty about a parameter.
2. Trace and rank plots
Inspect multiple chains for drift, persistent differences, long periods of
little movement and uneven exploration. Agreement across chains supports the
numerical assessment, but chains can agree while missing the same region.
Trace plots need not look like independent white noise: MCMC draws are
dependent by construction.
Use trace or rank plots alongside numerical diagnostics, including checks of
the parameters and derived quantities that drive the decision. A visually
stable trace alone is insufficient evidence of convergence.
3. R-hat: screen for disagreement
Modern R-hat uses split chains, rank normalisation and a folded comparison to
detect differences in location and scale. It improves on the original
between-chain/within-chain variance comparison. See
Vehtari et al. for the method and limitations.
Values close to one are desirable. Abacus’s default 1.01 threshold is a
screening convention, not a safe/unsafe boundary that proves convergence.
Values at or above the configured threshold are flagged. Non-finite or
unavailable diagnostics must be investigated rather than counted as passing.
When chains disagree, inspect their paths and the model geometry. More warmup
or retained draws may help, but persistent separation can require a different
parameterisation or a revised model. Simply extending a run is not a general
solution.
4. Effective sample size: Monte Carlo precision
Effective sample size (ESS) describes the precision of a simulation-based
summary relative to independent draws. It is not the number of observations
in the dataset or the model’s degrees of freedom. It is also not a Newey–West
standard-error correction: that procedure concerns estimation uncertainty
under dependent data, whereas MCMC ESS concerns dependent simulation draws.
Bulk ESS screens exploration of the main posterior mass; tail ESS helps assess
precision near interval endpoints. Adequacy depends on the quantity and
precision needed. Abacus’s default ESS threshold of 400 is a screening
convention, not a universal guarantee. Check Monte Carlo standard errors for
reported summaries; rare-event probabilities can need substantially more
simulation than a posterior mean.
If exploration is otherwise adequate, more retained draws can improve
precision. If chains mix poorly, investigate scaling and parameterisation.
Thinning an existing chain discards draws; it does not recover unexplored
regions and is not a general remedy for low ESS. Storage constraints are a
separate consideration.
5. Divergences: investigate retained transitions
A divergence signals excessive numerical error along a Hamiltonian
trajectory and can indicate posterior geometry that the sampler explores
poorly. It is evidence of a computational problem to investigate, not proof
of one particular model defect or a known amount of bias.
Separate warnings during warmup from divergences among retained draws.
Target zero retained divergences and investigate any that remain, even when
R-hat and ESS look acceptable. Do not dismiss a small retained count by calling
it warmup. Locate the affected parameter regions and examine sensitivity to
sampling settings and parameterisation. See the
Stan diagnostic guidance
for the computational interpretation.
Increasing target_accept can reduce integration error at additional
computational cost. Persistent divergences can require rescaling,
reparameterisation or revising the model. Adding draws alone does not repair
the underlying geometry. Zero observed divergences is useful evidence, but
not proof of adequate exploration or model validity.
Abacus distinguishes unavailable divergence evidence from an observed zero.
Missing or invalid retained flags produce an unavailable status and reason;
they must not be interpreted as a clean run. Consult the diagnostics guide
for the actual report fields and pipeline handling.
6. Interpret posterior intervals conditionally
A 95% credible interval contains 95% of the parameter’s posterior probability,
conditional on the data, likelihood and prior. That statement is different
from the repeated-sampling coverage of a confidence-interval procedure.
Neither interpretation removes the need to assess model assumptions.
A highest-density interval (HDI) and an equal-tailed interval are different
summaries and can differ for skewed distributions. State the interval method
and probability used; a probability label alone does not identify the method.
The interval calculation in Abacus summary facades is a separate API contract
from the interpretation of a posterior probability.
7. Interpret signs and practical thresholds
A 94% posterior interval above zero is not generally equivalent to rejecting
a null hypothesis at a 6% significance level. Posterior probabilities do not
provide that frequentist error-rate guarantee. See the
conditional interval interpretation.
First inspect the prior support. With a positive continuous prior such as a
HalfNormal, $P(\beta > 0 \mid y)=1$ by construction. An interval above zero does
not show that the data discovered positivity. Increasing the prior scale
cannot introduce negative support. See
Bayesian Priors.
Where the model permits the relevant alternatives, report an interval and, if
useful, posterior probability relative to a prespecified practical threshold.
For example, the probability that a response exceeds a meaningful minimum
addresses magnitude within the fitted model. Compare it with the prior
probability and assess prior sensitivity. Report probabilities estimated from
MCMC draws with appropriate Monte Carlo precision, not as exact calculations.
If an interval includes zero, the estimate is inconclusive as to sign at that
interval probability. If it rules out effects of practical importance, state
that narrower claim and the threshold. None of these summaries establishes
causal identification on its own.
8. Review the evidence before interpretation
Confirm that diagnostics describe retained draws and that the required
evidence is available. Missing evidence is not passing evidence.
Investigate retained divergences and flagged R-hat, bulk/tail ESS, energy
or tree-depth diagnostics. Inspect trace or rank plots for the same run.
Check Monte Carlo precision for the summaries that will be reported.
Increase sampling only where it addresses the diagnosed problem.
Separately assess predictive checks, prior sensitivity, model assumptions
and causal identification before using outputs for decisions.
Report the posterior summary, interval method and probability, computational
limitations and substantive assumptions. A computational screen can support
use of the draws for further analysis; it does not certify the conclusions.
Prior Predictive Checks
If you come from classical econometrics, you are used to checking assumptions
after estimation: residual plots, heteroskedasticity tests, outlier influence,
and maybe out-of-sample fit. Bayesian workflow adds one earlier question:
Before fitting anything, do my priors imply plausible behaviour for the
target variable?
That is what prior predictive checking answers.
1. Why parameter-level priors are not enough
A prior can look sensible when you inspect it in isolation and still imply
absurd behaviour once it flows through the whole model.
For example:
an intercept prior may look “weakly informative” on paper
a channel coefficient prior may look “reasonably positive”
a likelihood sigma prior may look “safely diffuse”
But jointly, those choices might imply:
weekly revenue that is far above anything you could ever observe
negative conversions for a business where the target is always non-negative
far more volatility than the real series could possibly have
Classical econometrics rarely forces you to check this explicitly because you
usually specify penalties or constraints directly on the coefficient space.
Bayesian MMM requires one more layer of discipline: inspect the implied
distribution of y, not just the configured priors on the parameters.
2. What prior predictive checking does
Prior predictive checking asks:
If the priors were true, what kinds of target series would this model
generate before seeing the actual data?
The workflow is:
Build the model with your chosen priors and structure.
Sample from the prior predictive distribution.
Compare those simulated target draws with the scale and shape of the real
target series.
This is not a convergence check and it is not a causal test. It is a
plausibility check on the model you are about to fit.
3. How Abacus supports it
Abacus exposes prior predictive sampling directly on PanelMMM:
In the structured runner, this is Stage 10, the preflight stage. The pipeline
writes:
10_pre_diagnostics/prior_predictive.nc
10_pre_diagnostics/prior_predictive.png
Abacus currently gives you the sampled draws and the plot. It does not
apply an automatic plausibility score or a hard pass/fail gate for you.
4. What to look for
A useful prior predictive check is not about matching the data exactly. That
would defeat the point of a prior. The question is whether the implied target
behaviour is at least in the right universe.
Look for the following.
Level
Do the simulated draws live on roughly the same order of magnitude as the
observed target?
If your historical weekly revenue is in the low millions, prior predictive
draws in the billions are a red flag.
Dispersion
Is the implied volatility remotely plausible?
If the prior predictive distribution is much wider than the observed series,
your likelihood sigma or contribution priors are probably too loose.
Sign and support
Does the model imply values that violate business reality?
For example:
negative conversions
implausibly negative revenue
large oscillations around zero for a strictly positive KPI
These are often signs that the prior scale is too permissive relative to the
data scaling and likelihood choice.
Time pattern
Do the implied trajectories look structurally plausible?
You are not looking for a perfect seasonal pattern before fitting, but you
should ask whether the prior predictive draws look like something that could
have come from your business rather than from a random-number generator with no
economic interpretation.
5. Common failure modes
Several practical pathologies show up repeatedly.
The intercept is too loose
A very wide intercept prior can dominate the prior predictive distribution,
especially when the target has been scaled but the intercept prior is still too
diffuse for the transformed space.
The likelihood sigma is too loose
If the prior predictive draws look far too noisy, the problem is often not the
media priors at all. It is the observation model allowing implausibly large
residual variance.
Media transformation priors are too permissive
Adstock and saturation priors that allow unrealistically persistent carryover
or unrealistically steep response can imply contributions that are wildly too
large before the data has had any say.
Flexible baseline terms are too unconstrained
Time-varying intercepts, seasonality, events, and other additive effects can
all inject structure into the prior predictive distribution. If those priors
are too loose, the target series can become implausibly volatile or
pattern-heavy before fitting.
6. What to do when the prior predictive check looks bad
Do not proceed directly to posterior interpretation. Fix the model first.
Typical remedies:
tighten the intercept prior
tighten the likelihood sigma prior
make media priors more weakly informative in the economically plausible
region rather than completely diffuse
reduce unnecessary model flexibility before the data has justified it
check whether your scaling choices make the configured priors too wide or too
narrow on the model scale
This is the Bayesian analogue of catching a broken specification before you
start arguing about p-values.
7. What prior predictive checks do not tell you
Passing a prior predictive check does not mean:
the model is causally identified
the model will fit well
the posteriors will converge cleanly
the attribution decomposition will be trustworthy
It only means the configured priors do not imply obviously absurd target
behaviour before seeing the data.
Treat prior predictive checking as a standard pre-fit step, not as an optional
extra for purists.
In Abacus terms, the workflow should usually be:
Specify the model and priors.
Run sample_prior_predictive(...).
Inspect the implied target behaviour.
Revise the priors if needed.
Fit only once the prior predictive behaviour is broadly plausible.
That sequence is usually cheaper than fitting a badly specified Bayesian MMM
and then discovering that the posterior is unstable for reasons you could have
caught before sampling.
Posterior Predictive Checks
Posterior predictive checking asks a simple question:
After fitting the model, can it reproduce the main features of the observed
data?
For a classically trained econometrician, this is the Bayesian analogue of
residual diagnostics, fitted-versus-observed checks, and out-of-sample
sanity-checking, but with one important difference: the checks are based on the
full posterior distribution, not a single point estimate.
1. What the check actually is
After fitting, you sample from the posterior predictive distribution:
That assessment stage is the closest Abacus comes to a retained,
systematically-produced posterior predictive diagnostics bundle.
4. What to inspect
Observed versus fitted over time
Start with the time-series overlay.
Ask:
Does the fitted mean track the major movements in the target?
Are the predictive intervals wide enough to cover the observed series
reasonably often?
Does the model systematically lag turning points or seasonal peaks?
If the observed line keeps sitting outside the predictive interval in
structured ways, the model is missing something systematic rather than merely
being noisy.
Residual structure
Residuals should not show strong unresolved patterns.
In practice, look for:
long runs of positive residuals followed by long runs of negative residuals
clear seasonality left in the residuals
residual variance increasing with fitted values
one panel slice fitting much worse than the others
The presence of structure in the residuals usually means the model is still
under-specified for the data.
Scatter of fitted versus observed
The fitted-versus-observed scatter is not a formal test, but it quickly shows:
compression toward the mean
systematic underprediction at high values
systematic overprediction at low values
This is the Bayesian cousin of the fitted-value plots you would inspect after a
classical regression.
5. What “good” posterior predictive behaviour looks like
A good posterior predictive check does not mean the model matches every
wiggle exactly.
You are looking for something more practical:
the main level and variation are captured
the observed series falls inside plausible predictive ranges often enough
residuals are not strongly structured
panel slices are not failing in obviously asymmetric ways
The question is whether the model is adequate for interpretation, not whether
it is perfect.
6. What posterior predictive checks cannot prove
This is the most important warning.
A model can pass posterior predictive checks and still fail as a causal model.
Why? Because posterior predictive checks evaluate prediction of the target,
not causal attribution of the components.
Two models can predict sales equally well while assigning very different shares
of those sales to:
baseline
media
controls
seasonality
events
That is why posterior predictive checking must be paired with:
If the fitted line misses broad movements or regime changes, the model may need
more structural flexibility, for example in trend, seasonality, controls, or
events.
The model is too flexible in the wrong place
You may see good in-sample fit but strange residual behaviour or unstable
attribution because the model is fitting noise through components that should
remain more constrained.
Media is carrying baseline structure
If media spend is strongly correlated with time patterns, the model may let
media soak up baseline variation that should have been handled by intercept,
seasonality, controls, or other additive structure.
Baseline is carrying media structure
The reverse can also happen: a very flexible baseline can absorb variation that
you would otherwise attribute to media.
8. What to do when checks fail
If posterior predictive checks look bad, resist the temptation to jump straight
to interpreting coefficients anyway.
Instead:
Check convergence first.
Inspect residual structure rather than only aggregate fit.
Revisit baseline specification, controls, seasonality, events, and media
transformation choices.
Refit and compare again.
In other words, use posterior predictive checking as a model-development tool,
not just as a reporting plot.
9. Practical recommendation
In Abacus, the robust sequence is:
Run prior predictive checks before fitting.
Fit the model and verify MCMC diagnostics.
Run posterior predictive checks and inspect residuals.
Only then move to contributions, optimisation, or causal interpretation.
That order mirrors how a careful econometrician would already work, except that
the Bayesian workflow makes the predictive-check step much richer and more
honest about uncertainty.
Model Comparison
Choose the prediction task before choosing a score. Abacus reports
leave-one-out (LOO) and WAIC diagnostics from the stored likelihood, but those
scores do not establish causal attribution or automatically assess forecasting.
What ELPD measures
For held-out units indexed by $i$, LOO estimates the expected log predictive
density (ELPD) using contributions of the form
$\log p(y_i \mid y_{-i})$. ArviZ reports their sum, elpd_loo, not their
average. Higher values indicate better predictive performance for the same
scoring task and measure. The absolute value depends on the target scale and
number of held-out units.
Pareto-smoothed importance sampling (PSIS) approximates these LOO calculations
from a fitted posterior, avoiding a refit for every unit when the approximation
is reliable. It still requires adequate posterior sampling and suitable
importance weights. For the Abacus accessors and report fields, see
Diagnostics.
Check comparability before ranking models
Require the same outcome definition and scale, likelihood measure, held-out
unit, observations and prediction task. Matching the input CSV is insufficient.
Abacus preset
Stored likelihood basis / held-out unit
time_series
Outcome level space, with pointwise contributions over dates
fe
Within-unit orthonormal contrasts; the outcome-level unit intercept is removed from this likelihood
cre
Marginal likelihood for each complete unit block, integrating over the random unit intercept
Use the estimator manifest and the relevant
estimator specification to
check the actual contract. The compact Bayesian-criteria report distinguishes
FE contrast_space from level_space; that field alone does not distinguish
CRE unit-block scoring from time-series date scoring.
For example, two time-series specifications fitted to the same dates and
outcome scale can be compared if both use the same likelihood measure and
acceptable LOO approximations. Comparing FE contrast-space ELPD directly with
CRE level-space ELPD is invalid even when both models use the same panel rows.
Their predictive densities score different quantities. Changing the target
from y to log(y) also requires reconciling the density measure before any
comparison; a label change is not sufficient.
Interpret differences and approximation diagnostics
az.compare(...) reports score differences and uncertainty based on the
pointwise contributions. Only supply models that satisfy the comparability
conditions above. Examine the size and practical relevance of the difference,
its estimated standard error and the observations driving it.
A difference small relative to its uncertainty is inconclusive as to which
model predicts better. It does not establish equivalence. A two-standard-error
rule is not a universal significance test, particularly with few held-out
units, dependence or influential observations. Simplicity and interpretability
can guide a decision when predictive evidence is inconclusive, but record that
as a decision criterion rather than a demonstrated equality of performance.
Check Pareto-k before relying on PSIS-LOO. The diagnostic threshold depends on
the number of draws, with 0.7 an upper cap on the usual threshold. Values from
0.7 to 1 indicate unreliable estimates with potentially substantial bias;
values at or above 1 indicate a more severe failure. Abacus’s fixed counts
above 0.7 and 1 are summaries, not a complete sample-size-dependent reliability
assessment. Inspect the ArviZ warnings and pointwise diagnostics too.
See the primary PSIS diagnostic reference.
For problematic observations, investigate the data and model, then consider
explicit refits or a suitable K-fold/blocked validation design. This guide does
not assume an Abacus or ArviZ moment-matching convenience API. Switching to
WAIC does not by itself resolve unreliable importance sampling or influential
observations.
Assess future prediction with held-out future data
Ordinary LOO can condition on observations later than the omitted date. That
is a different task from forecasting without future outcomes. For Abacus’s
separate training-prefix fit and later holdout window, use
Blocked Holdout Validation.
A single terminal block assesses that window; it is not a rolling-origin study
or a guarantee for every future horizon. The
leave-future-out case study
explains the distinction from ordinary LOO.
Keep prediction, model adequacy and attribution separate
Use posterior predictive checks
to inspect features relevant to the application, such as volatility and
residual time structure. Relative score improvements do not imply that either
model is adequate. Conversely, visually similar predictive fits can conceal
very different channel contributions.
AIC, BIC, Bayes factors and predictive scores address different objectives and
assumptions; they are not interchangeable replacements. In particular, plugging
a posterior mean into a likelihood and applying a nominal parameter penalty
does not automatically recover the usual AIC/BIC justification.
Assess attribution using the identifying assumptions, prior/baseline
sensitivity and relevant external evidence. LOO cannot establish causal
identification. See Causal Identification
and Baseline vs Media Trade-Offs.
Causal Identification
An Abacus fit estimates quantities conditional on the specified likelihood,
priors and data. Interpreting a media contrast as the effect of an intervention
requires a separate identification argument. Good predictive fit, stable
sampling and narrow posterior intervals do not supply that argument.
Define the intervention and identifying variation
State the channel, spending change, dates, population and outcome for the
causal question. Specify how carryover and changes in other channels enter
the contrast. Then explain why the observed variation identifies that effect.
For observational MMM, consequential assumptions include:
an adequate control set for common causes of media and outcomes, without
conditioning on mediators or colliders that invalidate the intended effect;
enough relevant variation to learn the response in the spending range of
interest, rather than relying entirely on extrapolation;
consistent treatment/outcome definitions, and appropriate treatment of
interference, carryover and anticipation;
an adequate response, baseline and error specification for the estimand.
For example, spending that responds to expected demand can be associated with
higher sales even without a media effect. Adding a smooth trend does not
necessarily account for that demand information. A control associated with
both variables is useful only if its causal role and measurement justify
conditioning on it. These assumptions require subject-matter evidence;
residual checks alone cannot establish them. See the primary discussion of
MMM causal assumptions.
The named FE estimator removes time-invariant unit effects from its slope
likelihood. CRE adjusts for its declared transformed between-unit summaries.
Neither removes arbitrary time-varying confounding. See
Choose an Estimator and
Baseline vs Media Trade-Offs.
Assess designs by their assumptions
There is no universal ranking of methods that substitutes for assessing the
design and estimand.
Design
Key distinction
Randomised experiment
Assignment supports an identification argument for the specified treatment contrast; adherence, missing outcomes, interference and the analysis population still matter
Matched-market study without random assignment
Matching does not create randomisation; identification depends on the design’s comparability and counterfactual assumptions
Instrumental variables
Requires relevance, independence and exclusion; some estimands also require monotonicity or further structural assumptions
Difference-in-differences
Requires an appropriate untreated-trend assumption and treatment-timing conditions; compatible pre-trends do not prove parallel counterfactual trends
Regression discontinuity
Identification near a cutoff relies on the assignment mechanism and continuity or local-randomisation assumptions; extrapolation requires more
Diagnostics can challenge aspects of a design. They do not generally prove
exclusion, absence of confounding or unobserved counterfactual behaviour.
Effects identified for different populations or interventions are not
interchangeable merely because each estimate is credible for its own task.
Use calibration for the evidence it supplies
Abacus exposes mmm.add_lift_test_measurements(...) and
mmm.add_cost_per_target_calibration(...) on a built model. Both are unavailable
for the named FE, CRE and release-gated RE presets. Follow
Calibration for supported inputs and
workflow. EventAdditiveEffect adds an event component; it does not attach
experimental lift evidence.
Before calibration, align the external estimate with the model’s channel,
units, spending contrast, time window and outcome. Assess the design’s
identification, uncertainty and relevance to the modelling population. Avoid
counting the same evidence twice without an appropriate dependence model.
Calibration conditions the fitted response on the supplied evidence under
the calibration model. Conflicting observations and calibration data need
investigation; their combination is not guaranteed to remove bias. A test for
one channel and window does not identify every other channel or establish
transportability to a different spending range. Record these limits alongside
the calibrated result.
Interpret scenarios and optimisation conditionally
A scenario evaluates a specified spending plan under the fitted response
functions and planning assumptions. Optimisation selects a plan for a chosen
objective and constraints. Neither operation independently verifies that the
intervention will produce the predicted change.
A channel ranking at historical spend is insufficient for allocation. The
optimum depends on marginal responses across feasible spending levels,
saturation, carryover and constraints. Two models can rank historical returns
in the same order yet recommend different allocations. Proportional bias in
some reported channel estimates does not establish that their full marginal
response functions are correct.
Report the estimand, units, assumptions, supported spending range and
uncertainty. Posterior intervals condition on the fitted model; they do not
automatically include uncertainty about omitted confounding, model choice or
future structural changes. Use sensitivity analyses and relevant experimental
evidence to assess those risks. See
Interpreting Optimisation.
Baseline vs Media Trade-offs
One of the most confusing experiences in MMM is this:
two specifications can fit the target series almost equally well
both can have acceptable diagnostics
yet they can assign very different amounts of the target to media versus
baseline
This is not necessarily a bug in the software. It is a structural feature of
the problem.
This page explains how that trade-off appears in Abacus and why you should
expect it.
1. The decomposition problem
At a high level, Abacus builds the expected target from several additive
components.
In the retained PanelMMM build path, the mean function can include:
intercept_contribution
channel_contribution
control_contribution, if you configure control_columns
mundlak_contribution, if use_mundlak_cre=True
yearly_seasonality_contribution, if yearly_seasonality is enabled
additional additive effects you attach before build, such as events or trend
effects
The likelihood sees the sum of these pieces, not a directly observed
“ground-truth baseline” and “ground-truth media” split.
That means the total fit can be easier to identify than the decomposition.
2. Why the trade-off exists
Suppose revenue rises every December and TV spend also rises every December.
Several stories can fit the same sales data reasonably well:
December uplift is mostly seasonality
December uplift is mostly TV
December uplift is partly both
If the model includes both a seasonal term and media terms, they will compete
to explain the same observed movement.
This is the core baseline-versus-media trade-off:
the data often identify total explained variation better than they identify
which component deserves the credit
Classical econometricians already know this as collinearity and omitted-variable
competition. Bayesian MMM does not make that problem disappear. It makes the
uncertainty around it more explicit.
3. What counts as “baseline” in Abacus
In Abacus, the baseline side comes from the terms you specify inside the PyMC
graph.
Depending on configuration, that can include:
a static intercept
a time-varying intercept
yearly Fourier seasonality
controls
events
trend-like additive effects
Mundlak CRE adjustments in panel settings
So when people say “baseline absorbed the effect”, they usually mean one or
more of those components, not a separate external decomposition engine.
4. How media can lose attribution
Media can lose attribution when the non-media side of the model is too good at
explaining the same movements.
Common cases:
a flexible time-varying intercept captures medium-run swings that media could
also explain
strong seasonal terms absorb repeating peaks that coincide with campaign
timing
control variables proxy for media timing or market conditions too strongly
event effects explain demand spikes that were previously being picked up by
channel coefficients
In each case, the model may still predict well. The question is how the
variation is partitioned.
5. How media can steal attribution from baseline
The reverse failure is also common.
If the baseline side is under-specified, media channels can absorb variation
that is not truly incremental media response.
Examples:
missing seasonality leaves recurring annual structure for media to explain
missing controls leave competitor, pricing, or macro effects for media to
explain
missing events leave spikes for channels to absorb
insufficient baseline flexibility forces media to act as a trend proxy
This usually inflates media contribution and makes optimisation outputs look
better than they should.
6. Why good fit does not settle the argument
You might hope that whichever specification predicts better must also have the
more trustworthy attribution split.
Unfortunately, that does not follow.
A model can reproduce the observed target series very well while still having
ambiguous attribution. Predictive adequacy is necessary, but it is not enough
to identify the correct media decomposition.
7. Signs that the trade-off is driving your result
Be cautious when you see any of the following:
very similar model fit with materially different channel contributions
large channel swings after adding or removing a seasonal or trend term
media ROI rankings that flip after adding controls or events
one highly flexible baseline term dominating decomposition while media
contributions collapse
implausibly smooth media contributions paired with a very wiggly baseline, or
vice versa
These are not proofs of misspecification, but they are strong prompts for
sensitivity analysis.
8. What to do in practice
A disciplined Abacus workflow is usually better than trying to argue
theoretically about the “right” split in the abstract.
Recommended approach:
Start with a specification that has the minimum baseline structure you can
defend.
Add seasonal, control, event, or time-varying terms only when you can
justify them substantively or diagnostically.
Refit and compare decomposition stability, not just target fit.
Report instability when attribution changes materially across defensible
specifications.
Where possible, bring in external evidence such as lift tests or
calibration.
The important point is not to force one narrative prematurely. It is to show
which attribution conclusions remain stable after reasonable specification
changes.
9. Abacus-specific interpretation
In Abacus, you should treat the decomposition outputs as conditional on the
configured structure:
the chosen controls
whether yearly_seasonality is on
whether the intercept is time-varying
whether media effects are time-varying
whether you added events or other additive effects
whether use_mundlak_cre=True
Change the structure, and the attribution can change even when predictive fit
does not move much.
That is normal. It is the software telling you where the data alone are not
decisive.
10. Bottom line
Baseline-versus-media trade-offs are unavoidable in MMM because the observed
target only reveals the sum of the contributing processes.
Abacus makes this explicit by fitting all configured terms inside one additive
Bayesian graph. That is a strength, but it also means you need to read the
decomposition as a conditional statement:
given this model structure, priors, and data, this is the most plausible
attribution split
That is much more defensible than pretending the split is uniquely observed in
the data.
Mundlak Specification Test
Background
Classical panel econometrics uses a Mundlak specification test to assess the
mean-independence restriction behind a random-effects (RE) model. In its usual
form, the test evaluates whether the coefficients on declared unit-level
regressor summaries are jointly zero.
A rejection is evidence against that particular RE restriction under the test
assumptions. A failure to reject is not proof that RE is adequate: weak
variation, collinearity, finite samples, or an incomplete summary basis can
leave the test uninformative.
Why Abacus Does Not Reproduce the Frequentist Test
Abacus fits Bayesian models. It does not attach an asymptotic Wald test or a
chi-squared reference distribution to the Mundlak coefficients. Posterior
inference on those coefficients answers a different question and depends on
the declared priors and summary basis.
Two interpretations must be avoided:
A posterior interval containing zero does not establish that the RE
mean-independence assumption is adequate.
A posterior interval excluding zero indicates a conditional association with
the declared summaries. It does not identify the amount of confounding or
establish causal identification.
This distinction is especially important in marketing mix modelling, where
media transformations are estimated and the available between-unit variation
may be weak.
Legacy Mundlak Surface and Named CRE Preset
use_mundlak_cre=True is the retained low-level panel surface. It adds legacy
Mundlak terms to an unlabelled dimensioned PanelMMM. It is not an alias for
the named cre estimator preset.
The named CRE preset has a separate, explicit contract. Its media summaries
use the declared transformed exposure basis and it records estimability
evidence. The released v1 surface retains explicit limits on prediction and
unsupported downstream operations.
Do not transfer an interpretation or diagnostic result from one surface to the
other without checking the actual fitted summary basis.
What to Inspect
Posterior summaries
For a legacy Mundlak fit, inspect the coefficients and their joint posterior
geometry:
Treat these summaries as evidence about associations conditional on the fitted
model. Check effective sample size, R-hat, posterior correlations, prior
sensitivity, and the amount of within- and between-unit variation before
interpreting their magnitude.
Estimability and sensitivity
Before relying on the adjustment:
Confirm that the declared summaries have non-zero between-unit variation.
Inspect rank, collinearity and condition-number diagnostics.
Compare posterior results under defensible prior alternatives.
Check that substantive conclusions are not driven by one summary-basis
choice.
Use prior and posterior predictive checks to detect implausible model
behaviour.
Predictive comparison
A predeclared held-out comparison can test whether adding the adjustment
improves prediction for the intended forecasting task. Use a split that
respects panel and temporal dependence. Predictive improvement does not by
itself validate the identifying assumption or convert observational
associations into causal effects.
Summary
Evidence
Supported conclusion
Unsupported conclusion
Adjustment interval includes zero
The data and prior do not clearly separate the coefficient from zero
RE is adequate
Adjustment interval excludes zero
Association with the declared unit-summary basis
Identified confounding or causal correction
Predictive comparison improves
Better prediction for the declared holdout task
Correct causal structure
Estimability diagnostics pass
No detected defect under the implemented screens
Global or posterior-wide identification
Abacus should therefore retain explicit posterior, estimability, sensitivity
and predictive evidence. It should not turn the Bayesian adjustment into a
binary RE-versus-CRE adequacy test.
References
Mundlak, Y. (1978). “On the Pooling of Time Series and Cross Section Data.”
Econometrica, 46(1), 69–85.
Vehtari, A., Gelman, A., & Gabry, J. (2017). “Practical Bayesian model
evaluation using leave-one-out cross-validation and WAIC.”
Statistics and Computing, 27(5), 1413–1432.