Cross-validation of species distribution models: removing spatial sorting bias and calibration with a null model

Species distribution models are usually evaluated with cross-validation. In this procedure evaluation statistics are computed from model predictions for sites of presence and absence that were not used to train (fit) the model. Using data for 226 species, from six regions, and two species distributi...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Ecology (Durham) 2012-03, Vol.93 (3), p.679-688
1. Verfasser:	Hijmans, Robert J
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Animal and plant ecology Animal, plant and microbial ecology Animals AUC Autocorrelation Bioclim Biological and medical sciences Calibration Climate Computer Simulation Conservation biology cross-validation Demography Ecological modeling Ecological niches Ecology Ecosystem Fundamental and applied biological sciences. Psychology General aspects Geographic regions MaxEnt model evaluation Model trains Modeling Models, Biological niche model pairwise distance sampling Sampling Sampling bias spatial autocorrelation Spatial models spatial sorting bias Species species distribution model Species Specificity
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Species distribution models are usually evaluated with cross-validation. In this procedure evaluation statistics are computed from model predictions for sites of presence and absence that were not used to train (fit) the model. Using data for 226 species, from six regions, and two species distribution modeling algorithms (Bioclim and MaxEnt), I show that this procedure is highly sensitive to "spatial sorting bias": the difference between the geographic distance from testing-presence to training-presence sites and the geographic distance from testing-absence (or testing-background) to training-presence sites. I propose the use of pairwise distance sampling to remove this bias, and the use of a null model that only considers the geographic distance to training sites to calibrate cross-validation results for remaining bias. Model evaluation results (AUC) were strongly inflated: the null model performed better than MaxEnt for 45% and better than Bioclim for 67% of the species. Spatial sorting bias and area under the receiver-operator curve (AUC) values increased when using partitioned presence data and random-absence data instead of independently obtained presence-absence testing data from systematic surveys. Pairwise distance sampling removed spatial sorting bias, yielding null models with an AUC close to 0.5, such that AUC was the same as null model calibrated AUC (cAUC). This adjustment strongly decreased AUC values and changed the ranking among species. Cross-validation results for different species are only comparable after removal of spatial sorting bias and/or calibration with an appropriate null model.
ISSN:	0012-9658 1939-9170
DOI:	10.1890/11-0826.1