Is my data suitable for Abacus, and what must I prepare?

Check two things before fitting: whether the data satisfy Abacus’s input contract, and whether they contain enough relevant variation to answer your question. A correctly formatted dataset can still support weak or ambiguous channel estimates.

How much history do I need?

There is no universal minimum number of weeks, markets or observations that makes an MMM reliable. Assess the history against the proposed model:

  • Does it cover the seasonal patterns and business regimes you want to model?
  • Do channels change independently enough to separate their responses?
  • Is there information about different exposure levels, including the range relevant to the proposed decision?
  • Is there enough earlier history for the chosen adstock horizon and enough data left after reserving a meaningful holdout?
  • For a panel estimator, is there usable variation within units and, where required, between units?

Three years of channels that always move together may contain less useful channel-specific information than a shorter period with distinct changes. Adding markets does not automatically add independent information either. Priors can regularise weakly informed parameters; they do not create the missing variation. See Bayesian Priors.

Should my media inputs be spend, impressions or something else?

Use a consistently measured exposure variable that matches the response you want to model. Abacus applies adstock and saturation to the channel columns you supply. It does not obtain a separate monetary spend series or convert impressions into money for you.

The choice also determines which downstream interpretation is valid:

Channel input Outcome Contribution divided by input means
Monetary spend Revenue in the same currency Model-conditional ROAS
Monetary spend Conversion count Conversions per unit of spend; its reciprocal is cost per conversion
Impressions Revenue Revenue per impression, not ROAS

The standard budget workflow treats allocated channel-input quantities as spend. Do not pass a monetary budget into a model trained on impressions and assume Abacus supplies the price conversion. That requires an explicit, reviewed mapping and a compatible planning workflow.

Keep currency, tax treatment, gross/net definitions and attribution dates consistent. A change in measurement can resemble a change in response.

What table must I provide?

For Python fitting, supply X as a DataFrame and y as a row-aligned Series or one-dimensional array. Keep the target out of X; include the date, channel, control and declared panel-dimension columns. A combined CSV is supported by the YAML and runner workflows, which separate the target for you.

For an aggregate time series, use one observation per date. For a panel, provide exactly one row for each expected date and panel-coordinate combination. Multiple panel dimensions require the declared rectangular grid; they do not represent arbitrary nested or incomplete groups.

Use a consistent time frequency and align the outcome, media and controls to the same periods. l_max=8 refers to eight input periods, including the current period, not necessarily eight weeks. Document aggregation rules: sums can make sense for spend and revenue, while a price or rate needs a justified aggregation method.

See Input Data Requirements and Panel Data Layout for the exact column and alignment contracts.

Are missing values the same as zero activity?

No. Abacus requires observed values for the channel, control and target cells you supply; it does not silently replace missing measurements with zero. Absent panel rows and duplicate rows also need resolution before fitting.

For example, consider two markets in one week:

Market Recorded TV spend Interpretation and action
North 0 Keep zero if the source confirms no TV spend occurred
South Missing Investigate the missing measurement; do not infer zero spend

If South’s sales are also missing, adding a row with zero sales invents an observation. Recover the measurement or choose and document a defensible treatment of the missing data. Any imputation introduces assumptions whose effect on the analysis should be assessed. If you change the analysis window or units, check the panel contract again.

What does Abacus preprocess automatically?

Abacus scales the target and channel inputs before constructing the model. It does not automatically scale controls, choose a control set, reconcile currencies, adjust for inflation or perform domain-aware missing-data repair.

Default target and channel scaling pools over the panel dimensions. It is not automatically separate scaling for each market. Check the configured scaling dimensions before choosing prior magnitudes or comparing parameters.

Supply the media inputs intended for the configured transforms. Do not pre-adstock or pre-saturate them and then apply the same transformations again inside Abacus. Check the outcome implications of your complete specification with prior predictive draws on the appropriate scale.

See Scaling and Preprocessing and Prior Predictive Checks.

Should I combine correlated channels or add more controls?

Neither action is an automatic remedy. If two channels always run together, separate contributions may depend heavily on priors and specification choices. Consider a combined channel when it represents a coherent exposure and the decision can be made at that combined level. Combining channels changes the estimand: it cannot then justify separate channel allocations. Changing media mix or unit costs can also make the combined response unstable.

If separate estimates are essential, seek additional identifying variation or relevant experimental evidence, and report the limits of the current data. Inspect transformed predictors as well as raw correlations; the model learns from adstocked and saturated exposures.

Choose controls from their role in the outcome and media-assignment process, not solely because they improve fit. A variable affected by media can block part of the effect you intended to estimate. More controls can also introduce collinearity. See Causal Identification.

What should I prepare before the first fit?

Record the target, units, time frequency, panel definition and decision question. Retain an audited input table, its source and all preprocessing decisions. Review variation, plausible controls, model scale and priors. Choose the estimator and reserve the validation window before comparing preferred results.

Use Blocked Holdout Validation for a training-prefix refit and later holdout. The supplied holdout media and controls make this a conditional prediction exercise; it does not establish that those inputs would have been known in a live forecast. Keep information from the holdout out of fitted preprocessing and model selection wherever that information would be unavailable at prediction time.

Continue with Which Abacus model should I choose?.