MCMC Diagnostics

Use MCMC diagnostics to assess whether retained simulation draws support the posterior summaries you want to report. They concern Monte Carlo exploration and precision. They do not establish that the model is correctly specified, that media effects are identified, or that predictions will generalise.

For Abacus methods, report fields, threshold comparisons and unavailable states, use the canonical diagnostics guide.

1. What the sampler does

Markov chain Monte Carlo (MCMC) approximates posterior expectations and probabilities using dependent simulation draws. Abacus’s usual NUTS workflow uses Hamiltonian Monte Carlo trajectories with an adaptive trajectory length. Warmup adapts sampling settings; retained draws are used for posterior summaries. A larger retained draw count does not by itself establish adequate exploration.

For example, two chains with 2,000 retained draws each provide 4,000 parameter vectors. Means, intervals and estimated posterior probabilities derived from those vectors have Monte Carlo error. Keep that simulation error distinct from posterior uncertainty about a parameter.

2. Trace and rank plots

Inspect multiple chains for drift, persistent differences, long periods of little movement and uneven exploration. Agreement across chains supports the numerical assessment, but chains can agree while missing the same region. Trace plots need not look like independent white noise: MCMC draws are dependent by construction.

Use trace or rank plots alongside numerical diagnostics, including checks of the parameters and derived quantities that drive the decision. A visually stable trace alone is insufficient evidence of convergence.

3. R-hat: screen for disagreement

Modern R-hat uses split chains, rank normalisation and a folded comparison to detect differences in location and scale. It improves on the original between-chain/within-chain variance comparison. See Vehtari et al. for the method and limitations.

Values close to one are desirable. Abacus’s default 1.01 threshold is a screening convention, not a safe/unsafe boundary that proves convergence. Values at or above the configured threshold are flagged. Non-finite or unavailable diagnostics must be investigated rather than counted as passing.

When chains disagree, inspect their paths and the model geometry. More warmup or retained draws may help, but persistent separation can require a different parameterisation or a revised model. Simply extending a run is not a general solution.

4. Effective sample size: Monte Carlo precision

Effective sample size (ESS) describes the precision of a simulation-based summary relative to independent draws. It is not the number of observations in the dataset or the model’s degrees of freedom. It is also not a Newey–West standard-error correction: that procedure concerns estimation uncertainty under dependent data, whereas MCMC ESS concerns dependent simulation draws.

Bulk ESS screens exploration of the main posterior mass; tail ESS helps assess precision near interval endpoints. Adequacy depends on the quantity and precision needed. Abacus’s default ESS threshold of 400 is a screening convention, not a universal guarantee. Check Monte Carlo standard errors for reported summaries; rare-event probabilities can need substantially more simulation than a posterior mean.

If exploration is otherwise adequate, more retained draws can improve precision. If chains mix poorly, investigate scaling and parameterisation. Thinning an existing chain discards draws; it does not recover unexplored regions and is not a general remedy for low ESS. Storage constraints are a separate consideration.

5. Divergences: investigate retained transitions

A divergence signals excessive numerical error along a Hamiltonian trajectory and can indicate posterior geometry that the sampler explores poorly. It is evidence of a computational problem to investigate, not proof of one particular model defect or a known amount of bias.

Separate warnings during warmup from divergences among retained draws. Target zero retained divergences and investigate any that remain, even when R-hat and ESS look acceptable. Do not dismiss a small retained count by calling it warmup. Locate the affected parameter regions and examine sensitivity to sampling settings and parameterisation. See the Stan diagnostic guidance for the computational interpretation.

Increasing target_accept can reduce integration error at additional computational cost. Persistent divergences can require rescaling, reparameterisation or revising the model. Adding draws alone does not repair the underlying geometry. Zero observed divergences is useful evidence, but not proof of adequate exploration or model validity.

Abacus distinguishes unavailable divergence evidence from an observed zero. Missing or invalid retained flags produce an unavailable status and reason; they must not be interpreted as a clean run. Consult the diagnostics guide for the actual report fields and pipeline handling.

6. Interpret posterior intervals conditionally

A 95% credible interval contains 95% of the parameter’s posterior probability, conditional on the data, likelihood and prior. That statement is different from the repeated-sampling coverage of a confidence-interval procedure. Neither interpretation removes the need to assess model assumptions.

A highest-density interval (HDI) and an equal-tailed interval are different summaries and can differ for skewed distributions. State the interval method and probability used; a probability label alone does not identify the method. The interval calculation in Abacus summary facades is a separate API contract from the interpretation of a posterior probability.

7. Interpret signs and practical thresholds

A 94% posterior interval above zero is not generally equivalent to rejecting a null hypothesis at a 6% significance level. Posterior probabilities do not provide that frequentist error-rate guarantee. See the conditional interval interpretation.

First inspect the prior support. With a positive continuous prior such as a HalfNormal, $P(\beta > 0 \mid y)=1$ by construction. An interval above zero does not show that the data discovered positivity. Increasing the prior scale cannot introduce negative support. See Bayesian Priors.

Where the model permits the relevant alternatives, report an interval and, if useful, posterior probability relative to a prespecified practical threshold. For example, the probability that a response exceeds a meaningful minimum addresses magnitude within the fitted model. Compare it with the prior probability and assess prior sensitivity. Report probabilities estimated from MCMC draws with appropriate Monte Carlo precision, not as exact calculations.

If an interval includes zero, the estimate is inconclusive as to sign at that interval probability. If it rules out effects of practical importance, state that narrower claim and the threshold. None of these summaries establishes causal identification on its own.

8. Review the evidence before interpretation

  1. Confirm that diagnostics describe retained draws and that the required evidence is available. Missing evidence is not passing evidence.
  2. Investigate retained divergences and flagged R-hat, bulk/tail ESS, energy or tree-depth diagnostics. Inspect trace or rank plots for the same run.
  3. Check Monte Carlo precision for the summaries that will be reported. Increase sampling only where it addresses the diagnosed problem.
  4. Separately assess predictive checks, prior sensitivity, model assumptions and causal identification before using outputs for decisions.

Report the posterior summary, interval method and probability, computational limitations and substantive assumptions. A computational screen can support use of the draws for further analysis; it does not certify the conclusions.