Abstract
Cross-validation (CV) has been widely used in GeoAI research to evaluate the performance of machine learning models. Often, a labeled data set is randomly split into training and validation data, and a machine learning model is trained on the training data and then evaluated on the validation data in an iterative manner. Such a random CV approach could lead to an overestimate of model performance on geographic data, due to the common existence of spatial autocorrelation. Random CV can generate many training and validation data instances that are spatially close, and a model trained on such training data can be considered as having already “peeked” into the nearby validation data, given their spatial closeness and likely attribute similarity. Spatial CV can help address this issue by splitting the data spatially rather than randomly, thereby increasing the independence between the training and validation data. While a number of spatial CV methods have been developed, they are scattered in the literature across multiple disciplines, including ecology, remote sensing, GIScience, and computer science. This chapter discusses four main spatial CV methods identified from the multidisciplinary literature, and uses two examples based on real-world data to demonstrate these methods in comparison with random CV.
| Original language | English |
|---|---|
| Title of host publication | Handbook of Geospatial Artificial Intelligence |
| Publisher | CRC Press |
| Pages | 201-214 |
| Number of pages | 14 |
| ISBN (Electronic) | 9781003814924 |
| ISBN (Print) | 9781032311661 |
| DOIs | |
| State | Published - Jan 1 2023 |
Fingerprint
Dive into the research topics of 'Spatial Cross-Validation for GeoAI'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver