Detecting virtual concept drift of regressors without ground truth values

Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications the true value of the outcome variable we want to predict is unknown outside the training data, i.e., the ground truth is u...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Data mining and knowledge discovery 2021-05, Vol.35 (3), p.726-747
Hauptverfasser:	Oikarinen, Emilia, Tiittanen, Henri, Henelius, Andreas, Puolamäki, Kai
Format:	Artikel
Sprache:	eng
Schlagworte:	Artificial Intelligence Chemistry and Earth Sciences Computer Science Data Mining and Knowledge Discovery Drift Estimates Information Storage and Retrieval Machine learning Physics Real variables Regression analysis Special Issue of the Journal Track of ECML PKDD 2021 Statistics for Engineering
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications the true value of the outcome variable we want to predict is unknown outside the training data, i.e., the ground truth is unknown. Phenomena such as overfitting and concept drift make it difficult to directly observe when the estimate from a model potentially is wrong. In this paper we present an efficient framework for estimating the generalization error of regression functions, applicable to any family of regression functions when the ground truth is unknown. We present a theoretical derivation of the framework and empirically evaluate its strengths and limitations. We find that it performs robustly and is useful for detecting concept drift in datasets in several real-world domains.
ISSN:	1384-5810 1573-756X
DOI:	10.1007/s10618-021-00739-7