A calibration hierarchy for risk models was defined: from utopia to empirical data

Abstract Objective Calibrated risk models are vital for valid decision support. We define four levels of calibration and describe implications for model development and external validation of predictions. Study Design and Setting We present results based on simulated data sets. Results A common defi...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Journal of clinical epidemiology 2016-06, Vol.74, p.167-176
Hauptverfasser:	Van Calster, Ben, Nieboer, Daan, Vergouwe, Yvonne, De Cock, Bavo, Pencina, Michael J, Steyerberg, Ewout W
Format:	Artikel
Sprache:	eng
Schlagworte:	Bias Calibration Comparative analysis Computer Simulation Decision curve analysis Decision making Decision Support Techniques Epidemiology External validation Humans Internal Medicine Loess Logistics Models, Statistical Overfitting Reproducibility of Results Risk Risk Assessment Risk prediction models
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Abstract Objective Calibrated risk models are vital for valid decision support. We define four levels of calibration and describe implications for model development and external validation of predictions. Study Design and Setting We present results based on simulated data sets. Results A common definition of calibration is “having an event rate of R % among patients with a predicted risk of R %,” which we refer to as “moderate calibration.” Weaker forms of calibration only require the average predicted risk (mean calibration) or the average prediction effects (weak calibration) to be correct. “Strong calibration” requires that the event rate equals the predicted risk for every covariate pattern. This implies that the model is fully correct for the validation setting. We argue that this is unrealistic: the model type may be incorrect, the linear predictor is only asymptotically unbiased, and all nonlinear and interaction effects should be correctly modeled. In addition, we prove that moderate calibration guarantees nonharmful decision making. Finally, results indicate that a flexible assessment of calibration in small validation data sets is problematic. Conclusion Strong calibration is desirable for individualized decision support but unrealistic and counter productive by stimulating the development of overly complex models. Model development and external validation should focus on moderate calibration.
ISSN:	0895-4356 1878-5921
DOI:	10.1016/j.jclinepi.2015.12.005