Bayesian Priors

Priors constrain and regularise the model. Their influence depends on their support and scale, the likelihood and the information in the data. A proper posterior or stable fit does not by itself establish identification.

Support restrictions and prior scale

Separate two choices:

  • Support: which parameter values the model permits.
  • Concentration: how it distributes probability over those values.

A HalfNormal prior gives zero probability to negative values. A LogNormal prior restricts its parameter to strictly positive values. Increasing either prior’s scale does not allow the posterior to become negative: the likelihood cannot create posterior support where the prior assigns none.

Consequently, P(beta > 0 | data) is one by construction for a coefficient with a positive continuous prior. It is not evidence that the data established a positive effect. An interval above zero must be interpreted in light of that restriction. A posterior can still concentrate near zero; whether it rules out a practically negligible effect is a separate question.

A concentrated prior can reduce variance while introducing bias when its assumptions are wrong. A diffuse prior can permit implausible values or weakly identified combinations. Neither choice guarantees accurate estimates, good sampling or causal validity. “Weakly informative” only has meaning relative to the parameterisation and input/output scales.

Relation to penalised estimation

Under a specified likelihood, a Normal prior yields a quadratic penalty in the negative log posterior, and a Laplace prior yields an absolute-value penalty. This connects posterior modes to ridge and lasso estimates in the corresponding models. It does not make their full uncertainty calculations identical.

An unconstrained frequentist estimate does not require a Bayesian prior. A flat density on the whole real line is improper, and flatness changes under non-linear reparameterisation. Avoid treating it as an assumption-free probability distribution.

Specify priors in Abacus

Use Prior from pymc_extras.prior. For example, these transform objects restrict the response amplitude to be non-negative and the decay to (0, 1):

from pymc_extras.prior import Prior

from abacus.mmm import GeometricAdstock, LogisticSaturation

adstock = GeometricAdstock(
    l_max=4,
    priors={"alpha": Prior("Beta", alpha=1, beta=3)},
)
saturation = LogisticSaturation(
    priors={
        "beta": Prior("HalfNormal", sigma=2),
        "lam": Prior("Gamma", alpha=3, beta=1),
    },
)

This is a configuration fragment: pass the objects as adstock and saturation when constructing PanelMMM. The numerical scales above are illustrative, not a recommendation for every dataset. Abacus applies media transforms to scaled inputs, so assess their implications on the outcome scale too. For model-level overrides and valid YAML syntax, use Priors and Configuration.

Choose and assess a prior specification

  1. State the support restrictions and their substantive justification.
  2. Check plausible parameter magnitudes in the actual model scales.
  3. Inspect prior predictive draws for plausible outcomes before interpreting a fit.
  4. Assess the available identifying variation, including correlated media, persistent unit differences and possible time-varying confounding.
  5. Fit defensible alternative prior specifications and compare the quantities used for decisions, including contributions and scenario contrasts.

There is no universal threshold in weeks or number of channels that makes media effects identified or a prior negligible. More observations need not supply independent variation. External calibration can inform a particular response, but its relevance depends on the experimental design, estimand and transport assumptions; see Calibration.

Compare like quantities

Compare each parameter prior with its parameter posterior to inspect updating. Similar distributions do not uniquely diagnose an excessively concentrated prior: the likelihood may be weak, compatible with the prior, or informative about combinations rather than individual parameters. A shifted posterior also does not show that the prior has ceased to matter.

Compare prior predictive and posterior predictive distributions with observed outcomes to assess their implications for data. These distributions include different sources of uncertainty from a parameter distribution and cannot be substituted for it. Sensitivity analysis is needed to assess how conclusions depend on the prior, even when the posterior looks concentrated.

For the broader distinction between fit and attribution, see Baseline vs Media Trade-Offs and Causal Identification.