Model Comparison

Choose the prediction task before choosing a score. Abacus reports leave-one-out (LOO) and WAIC diagnostics from the stored likelihood, but those scores do not establish causal attribution or automatically assess forecasting.

What ELPD measures

For held-out units indexed by $i$, LOO estimates the expected log predictive density (ELPD) using contributions of the form $\log p(y_i \mid y_{-i})$. ArviZ reports their sum, elpd_loo, not their average. Higher values indicate better predictive performance for the same scoring task and measure. The absolute value depends on the target scale and number of held-out units.

Pareto-smoothed importance sampling (PSIS) approximates these LOO calculations from a fitted posterior, avoiding a refit for every unit when the approximation is reliable. It still requires adequate posterior sampling and suitable importance weights. For the Abacus accessors and report fields, see Diagnostics.

Check comparability before ranking models

Require the same outcome definition and scale, likelihood measure, held-out unit, observations and prediction task. Matching the input CSV is insufficient.

Abacus preset Stored likelihood basis / held-out unit
time_series Outcome level space, with pointwise contributions over dates
fe Within-unit orthonormal contrasts; the outcome-level unit intercept is removed from this likelihood
cre Marginal likelihood for each complete unit block, integrating over the random unit intercept

Use the estimator manifest and the relevant estimator specification to check the actual contract. The compact Bayesian-criteria report distinguishes FE contrast_space from level_space; that field alone does not distinguish CRE unit-block scoring from time-series date scoring.

For example, two time-series specifications fitted to the same dates and outcome scale can be compared if both use the same likelihood measure and acceptable LOO approximations. Comparing FE contrast-space ELPD directly with CRE level-space ELPD is invalid even when both models use the same panel rows. Their predictive densities score different quantities. Changing the target from y to log(y) also requires reconciling the density measure before any comparison; a label change is not sufficient.

Interpret differences and approximation diagnostics

az.compare(...) reports score differences and uncertainty based on the pointwise contributions. Only supply models that satisfy the comparability conditions above. Examine the size and practical relevance of the difference, its estimated standard error and the observations driving it.

A difference small relative to its uncertainty is inconclusive as to which model predicts better. It does not establish equivalence. A two-standard-error rule is not a universal significance test, particularly with few held-out units, dependence or influential observations. Simplicity and interpretability can guide a decision when predictive evidence is inconclusive, but record that as a decision criterion rather than a demonstrated equality of performance.

Check Pareto-k before relying on PSIS-LOO. The diagnostic threshold depends on the number of draws, with 0.7 an upper cap on the usual threshold. Values from 0.7 to 1 indicate unreliable estimates with potentially substantial bias; values at or above 1 indicate a more severe failure. Abacus’s fixed counts above 0.7 and 1 are summaries, not a complete sample-size-dependent reliability assessment. Inspect the ArviZ warnings and pointwise diagnostics too. See the primary PSIS diagnostic reference.

For problematic observations, investigate the data and model, then consider explicit refits or a suitable K-fold/blocked validation design. This guide does not assume an Abacus or ArviZ moment-matching convenience API. Switching to WAIC does not by itself resolve unreliable importance sampling or influential observations.

Assess future prediction with held-out future data

Ordinary LOO can condition on observations later than the omitted date. That is a different task from forecasting without future outcomes. For Abacus’s separate training-prefix fit and later holdout window, use Blocked Holdout Validation. A single terminal block assesses that window; it is not a rolling-origin study or a guarantee for every future horizon. The leave-future-out case study explains the distinction from ordinary LOO.

Keep prediction, model adequacy and attribution separate

Use posterior predictive checks to inspect features relevant to the application, such as volatility and residual time structure. Relative score improvements do not imply that either model is adequate. Conversely, visually similar predictive fits can conceal very different channel contributions.

AIC, BIC, Bayes factors and predictive scores address different objectives and assumptions; they are not interchangeable replacements. In particular, plugging a posterior mean into a likelihood and applying a nominal parameter penalty does not automatically recover the usual AIC/BIC justification.

Assess attribution using the identifying assumptions, prior/baseline sensitivity and relevant external evidence. LOO cannot establish causal identification. See Causal Identification and Baseline vs Media Trade-Offs.