FAQ · Forecast calibration

What is forecast calibration?

Reviewed against the platform's code on Sep 26, 2026

Forecast calibration is the agreement between the probabilities a forecast states and how often things actually happen. A calibrated forecaster's 70% calls come true about 70% of the time, and its 90% prediction intervals contain about 90% of outcomes. It should be measured on outcomes the model has not seen, and it is a different property from accuracy.

Why it matters

A probability is only useful if it means what it says. Position sizes, risk limits and stop distances all take a forecast's stated uncertainty at face value: a 90% interval that contains only 75% of outcomes understates risk, and positions sized on it are too large. Calibration alone is not enough, though. A forecaster who always predicts the long-run average rate is well calibrated and says nothing about any particular day. The standard goal, set out by Gneiting, Balabdaoui and Raftery (2007), is to make forecasts as sharp as possible subject to calibration: intervals as narrow as they can be while still holding their stated share of outcomes.

How it works

For forecasts of events, a reliability diagram sorts forecasts into bins such as 0–10% and 10–20%, and plots each bin's average forecast against how often the event happened; a calibrated forecaster sits on the diagonal. Expected calibration error condenses the diagram into one number: the gap between the two in each bin, averaged with weights by the number of forecasts in the bin. The Brier score is the mean squared difference between forecast probabilities and the 0 or 1 outcomes (Brier, 1950); it rewards calibration and resolution together. For intervals, calibration means coverage: about 90% of outcomes should fall inside a 90% band. Isotonic regression and Platt scaling recalibrate probabilities on held-out data; conformal prediction does the same for intervals.

How Opulence Alpha applies it

Opulence Alpha checks calibration rather than assuming it. Each trading day its combined stock forecasts carry 90% intervals at 1, 5, 21 and 63 trading days, and the day is stamped “ok”, “degraded” or “uncalibrated”. At publication it is ok only if coverage is within five percentage points of 90% and interval widths differ from stock to stock; failing either means degraded, and too little data means uncalibrated, not a pass. That first check can only use calibration data, so the stamp is revised on out-of-sample coverage once at least 100 outcomes have matured; the forecasts stay as published. The Regime Radar's five-session chance of a regime change is calibrated by isotonic regression followed by Platt scaling and shown with a 95% bootstrap interval.

Today's calibrated five-session regime forecast →

Questions

What is the difference between calibration and accuracy?

Accuracy is how close forecasts land to outcomes; calibration is whether their stated probabilities can be taken at face value. A forecaster that always predicts the long-run base rate is well calibrated yet useless on any given day, and a model that ranks events well can still say 90% for events that happen 60% of the time. Good forecasts need both, which is why proper scoring rules such as the Brier score reward calibration and resolution together.

How do you read a reliability diagram?

Each point is a bin of forecasts: its average forecast probability on the horizontal axis, and how often the event then happened on the vertical. Points on the diagonal are calibrated. Below it, the event happened less often than forecast; above it, more often. A curve flatter than the diagonal is the usual sign of overconfidence, with high probabilities too high and low ones too low. Check how many forecasts sit in each bin, because a thin bin can stray from the diagonal by chance.

How is forecast calibration tested in finance?

Mostly as coverage. A 99% value-at-risk figure should be exceeded on about 1% of days, and a 90% return interval should miss about one outcome in ten. Christoffersen (1998) tests two things: whether the share of misses matches the stated level, and whether misses arrive independently rather than bunching together in stressed periods. For full return distributions, Diebold, Gunther and Tay (1998) check where each outcome falls within its own forecast distribution; if the forecasts are right, those positions are uniformly spread and independent over time.

Is every percentage on Opulence Alpha's Regime Radar a calibrated probability?

No, and the page says which is which. The chance that the composite regime changes within five sessions is a calibrated probability, shown with its interval. The early-warning level beside it blends four warning signals into a score between 0 and 1, and the page labels it a composite, not a probability. A score that ranks risk well is not a probability until it has been calibrated against outcomes.

References

  • Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.
  • Christoffersen, P. F. (1998). Evaluating Interval Forecasts. International Economic Review, 39(4), 841–862.
  • Diebold, F. X., Gunther, T. A. & Tay, A. S. (1998). Evaluating Density Forecasts with Applications to Financial Risk Management. International Economic Review, 39(4), 863–883.
  • Gneiting, T., Balabdaoui, F. & Raftery, A. E. (2007). Probabilistic Forecasts, Calibration and Sharpness. Journal of the Royal Statistical Society: Series B, 69(2), 243–268.

Educational content about research methods. Not investment advice.