Publication:

Learn Here to Predict There: Optimizing Training Region Selection for Geographically Generalizable Wildfire Prediction

Loading...
Thumbnail Image

Files

written_final_report.pdf (23.07 MB)

Date

2026-04-16

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

In this paper, I investigate how to optimally select training regions for wildfire prediction models that generalize effectively across the diverse climates of the contiguous United States. Using the U.S. Department of Agriculture's Spatial Wildfire Occurrence Data for the United States, I begin by addressing discrepancies in reporting across states by applying a regression discontinuity design, in order to determine a minimum fire size threshold that minimizes the existence of discontinuities while still retaining a sufficiently large dataset. Then, I use a k-means-based approach to construct 352 geographically-localized clusters of observed fires. In order to identify which clusters are optimal for training, I propose a greedy selection algorithm that chooses clusters maximizing predictive performance across all clusters under three different weighting schemes: unweighted, fire-size weighted, and population-weighted. I find that the proposed selection method consistently outperforms random baselines in predictive accuracy. This research thus provides an optimal method of selecting regions for future wildfire-related data collection and research, yielding statistically significant improvements in the generalizability of predictive models compared to randomly selected locations.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation