Confidence intervals versus observed prediction error

Each point below is a held-out observation. The horizontal coordinate is the absolute residual, \(|y_i - \hat y_i|\); the vertical coordinate is the approximate 95% confidence-interval half-width, \(1.96\sqrt{\max(\widehat{V}_{IJ,i}, 0)}\). All models use 2,000 estimators and calibrate=False. Negative raw IJ variances are clipped only to compute the square root; each caption reports their number. A zero width caused by clipping is not evidence of certainty.

The dashed diagonal marks equal magnitudes. Points above it have observed errors smaller than the estimated half-width; points below it have larger errors. This is a diagnostic comparison, not a prediction-interval coverage validation: IJ uncertainty describes variation of the fitted prediction under training-set resampling. Observed residuals also contain irreducible outcome noise and model bias, so a nominal 95% confidence interval need not contain 95% of individual outcomes. The observed residual is not the unknown error relative to the true conditional mean.

For regression, the prediction is the mean of the individual estimator predictions. For classification, we use predict_proba(X_test)[:, k] and compute its IJ variance with class_index=k. The observed residual is \(|\mathbf{1}(y_i = c_k) - \hat p_k(x_i)|\), where \(c_k\) is forest.classes_[k]. Binary tasks use column 1; Wine has a separate panel for each of its three classes. The classes are not treated as independent, and these marginal intervals do not provide simultaneous coverage. The normal approximation is descriptive and is not clipped to [0, 1].

The first nine panels use the same seven datasets, 80/20 splits, and reference forests as Empirical Bayes Calibration Benchmark. The three Wine panels also match the multiclass gallery example. Auto MPG covers the random-forest regression gallery example, using the benchmark’s 20% test split rather than the gallery’s 25%. Additional panels cover the spam classifier (20% test split) and Auto MPG bagged SVR (25% test split), with their gallery model settings except that the ensemble size is increased to 2,000. All data-generation, split and model seeds are 42. California Housing uses all 20,640 rows.

Every documentation build executes examples/generate_calibration_benchmark.py to refit the models and regenerate these figures and the calibration table. No saved benchmark results or plot images are used as build inputs.

Auto MPG

Observed absolute error versus estimated CI half-width for Auto MPG.

79 held-out samples; 0 negative raw IJ variances clipped to zero for plotting.

California

Observed absolute error versus estimated CI half-width for California.

4,128 held-out samples; 1300 negative raw IJ variances clipped to zero for plotting.

Diabetes

Observed absolute error versus estimated CI half-width for Diabetes.

89 held-out samples; 0 negative raw IJ variances clipped to zero for plotting.

Breast Cancer class 1

Observed absolute error versus estimated CI half-width for Breast Cancer class 1.

114 held-out samples; 29 negative raw IJ variances clipped to zero for plotting.

Synth Hard class 1

Observed absolute error versus estimated CI half-width for Synth Hard class 1.

400 held-out samples; 156 negative raw IJ variances clipped to zero for plotting.

Synthetic Reg

Observed absolute error versus estimated CI half-width for Synthetic Reg.

200 held-out samples; 1 negative raw IJ variances clipped to zero for plotting.

Wine class 0

Observed absolute error versus estimated CI half-width for Wine class 0.

36 held-out samples; 12 negative raw IJ variances clipped to zero for plotting.

Wine class 1

Observed absolute error versus estimated CI half-width for Wine class 1.

36 held-out samples; 4 negative raw IJ variances clipped to zero for plotting.

Wine class 2

Observed absolute error versus estimated CI half-width for Wine class 2.

36 held-out samples; 13 negative raw IJ variances clipped to zero for plotting.

Spam

Observed absolute error versus estimated CI half-width for Spam.

1,000 held-out samples; 422 negative raw IJ variances clipped to zero for plotting.

Auto MPG bagged SVR

Observed absolute error versus estimated CI half-width for Auto MPG bagged SVR.

98 held-out samples; 0 negative raw IJ variances clipped to zero for plotting.