Publication: Learn Here to Predict There: Optimizing Training Region Selection for Geographically Generalizable Wildfire Prediction
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
In this paper, I investigate how to optimally select training regions for wildfire prediction models that generalize effectively across the diverse climates of the contiguous United States. Using the U.S. Department of Agriculture's Spatial Wildfire Occurrence Data for the United States, I begin by addressing discrepancies in reporting across states by applying a regression discontinuity design, in order to determine a minimum fire size threshold that minimizes the existence of discontinuities while still retaining a sufficiently large dataset. Then, I use a k-means-based approach to construct 352 geographically-localized clusters of observed fires. In order to identify which clusters are optimal for training, I propose a greedy selection algorithm that chooses clusters maximizing predictive performance across all clusters under three different weighting schemes: unweighted, fire-size weighted, and population-weighted. I find that the proposed selection method consistently outperforms random baselines in predictive accuracy. This research thus provides an optimal method of selecting regions for future wildfire-related data collection and research, yielding statistically significant improvements in the generalizability of predictive models compared to randomly selected locations.