Model Comparison
Choose the prediction task before choosing a score. Abacus reports leave-one-out (LOO) and WAIC diagnostics from the stored likelihood, but those scores do not establish causal attribution or automatically assess forecasting.
What ELPD measures
For held-out units indexed by $i$, LOO estimates the expected log predictive
density (ELPD) using contributions of the form
$\log p(y_i \mid y_{-i})$. ArviZ reports their sum, elpd_loo, not their
average. Higher values indicate better predictive performance for the same
scoring task and measure. The absolute value depends on the target scale and
number of held-out units.
Pareto-smoothed importance sampling (PSIS) approximates these LOO calculations from a fitted posterior, avoiding a refit for every unit when the approximation is reliable. It still requires adequate posterior sampling and suitable importance weights. For the Abacus accessors and report fields, see Diagnostics.
Check comparability before ranking models
Require the same outcome definition and scale, likelihood measure, held-out unit, observations and prediction task. Matching the input CSV is insufficient.
| Abacus preset | Stored likelihood basis / held-out unit |
|---|---|
time_series |
Outcome level space, with pointwise contributions over dates |
fe |
Within-unit orthonormal contrasts; the outcome-level unit intercept is removed from this likelihood |
cre |
Marginal likelihood for each complete unit block, integrating over the random unit intercept |
Use the estimator manifest and the relevant
estimator specification to
check the actual contract. The compact Bayesian-criteria report distinguishes
FE contrast_space from level_space; that field alone does not distinguish
CRE unit-block scoring from time-series date scoring.
For example, two time-series specifications fitted to the same dates and
outcome scale can be compared if both use the same likelihood measure and
acceptable LOO approximations. Comparing FE contrast-space ELPD directly with
CRE level-space ELPD is invalid even when both models use the same panel rows.
Their predictive densities score different quantities. Changing the target
from y to log(y) also requires reconciling the density measure before any
comparison; a label change is not sufficient.
Interpret differences and approximation diagnostics
az.compare(...) reports score differences and uncertainty based on the
pointwise contributions. Only supply models that satisfy the comparability
conditions above. Examine the size and practical relevance of the difference,
its estimated standard error and the observations driving it.
A difference small relative to its uncertainty is inconclusive as to which model predicts better. It does not establish equivalence. A two-standard-error rule is not a universal significance test, particularly with few held-out units, dependence or influential observations. Simplicity and interpretability can guide a decision when predictive evidence is inconclusive, but record that as a decision criterion rather than a demonstrated equality of performance.
Check Pareto-k before relying on PSIS-LOO. The diagnostic threshold depends on the number of draws, with 0.7 an upper cap on the usual threshold. Values from 0.7 to 1 indicate unreliable estimates with potentially substantial bias; values at or above 1 indicate a more severe failure. Abacus’s fixed counts above 0.7 and 1 are summaries, not a complete sample-size-dependent reliability assessment. Inspect the ArviZ warnings and pointwise diagnostics too. See the primary PSIS diagnostic reference.
For problematic observations, investigate the data and model, then consider explicit refits or a suitable K-fold/blocked validation design. This guide does not assume an Abacus or ArviZ moment-matching convenience API. Switching to WAIC does not by itself resolve unreliable importance sampling or influential observations.
Assess future prediction with held-out future data
Ordinary LOO can condition on observations later than the omitted date. That is a different task from forecasting without future outcomes. For Abacus’s separate training-prefix fit and later holdout window, use Blocked Holdout Validation. A single terminal block assesses that window; it is not a rolling-origin study or a guarantee for every future horizon. The leave-future-out case study explains the distinction from ordinary LOO.
Keep prediction, model adequacy and attribution separate
Use posterior predictive checks to inspect features relevant to the application, such as volatility and residual time structure. Relative score improvements do not imply that either model is adequate. Conversely, visually similar predictive fits can conceal very different channel contributions.
AIC, BIC, Bayes factors and predictive scores address different objectives and assumptions; they are not interchangeable replacements. In particular, plugging a posterior mean into a likelihood and applying a nominal parameter penalty does not automatically recover the usual AIC/BIC justification.
Assess attribution using the identifying assumptions, prior/baseline sensitivity and relevant external evidence. LOO cannot establish causal identification. See Causal Identification and Baseline vs Media Trade-Offs.