paper

Racing to the Starting Line: Measuring Opportunity and Improvement in Cross Country

arXiv:2509.10600

Abstract

Collegiate cross country programs set race schedules without large comparable datasets across courses and conditions. Using the comprehensive-era National Running Club Database (NRCD; 23,360 results; 7,056 athletes; 2023-2025), where course and weather coverage exceed 99%, we ask what public results can identify about seasonal improvement, measurement, and racing opportunity. Hierarchical models of Standardized times recover pooled within-season improvement of about -5.7 s/week for men and -5.2 for women, but reading those slopes as fitness requires residual meet difficulty not to track the calendar: crossed meet effects can erase the week slope, while repeating-venue effects recover a clearer men's estimate under a weaker assumption. After the opener, race-1-only forecasts of who will improve are near noise (held-out R^2 ~ 0). Weather standardization remains a useful convention (Converted Only overstates first-to-last gains by about 15-21 s relative to Standardized), and prospectively trained venue factors modestly help where prior course data exist. Successful programs race more and place higher at nationals, but that may not be why they win: the schedule-placement link is between programs and is entangled with roster size, not evidence that changing one team's race count raises nationals place.