Built with Hetarth Goyani as a two-person project. The question was whether online signal — what people post, what gets reported — adds anything to the federal risk data already available for a county.
Fusing five sources
1,020 wildfire events across 125 counties in California, Oregon, Washington, Arizona and Nevada, spanning 1962 to 2023. Ground truth came from SHELDUS hazard-loss records; the predictors came from the FEMA National Risk Index (467 features per county), state climate data, Reddit posts and online news coverage.
Damage became a three-class label — low, medium, high — cut at the 33rd and 67th percentiles of inflation-adjusted dollar losses.
Getting text onto a map
Attributing an unstructured Reddit post to a county is the part nobody advertises. It ran as a tiered pipeline: geo-tags first where they exist, then subreddit heuristics, then Claude-based place extraction checked against a county gazetteer, then a contextual and bounding-box fallback.
Feature engineering pushed the table to roughly 475–500 columns — risk aggregations, expected-annual-loss rollups, vulnerability-to-resilience ratios, climate interaction terms — then mutual-information and model-based selection cut it back to the top 300–350.
Results, and the honest caveat
Eight classifier families were tuned with RandomizedSearchCV and Optuna, with SMOTE handling the class imbalance, then combined into soft-voting, stacking and weighted-voting ensembles.
The top predictors were the FEMA wildfire risk metrics, the social-vulnerability to resilience ratio, and raw social and news volume.
The caveat is the interesting result. Where the model misses, it misses in counties with thin Reddit and news coverage — data blackspots. Those were modelled explicitly, with missingness indicators and a deliberate choice between zero- and NA-imputation, rather than letting sparse coverage quietly bias the prediction toward whatever the federal data alone implies.
