Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Uncovering Insights of Compound Flooding with Data-Driven AI
Xu Zheng, Chaohao Lin, Sipeng Chen +7
Compound flooding, driven by nonlinear interactions between multiple hydrometeorological factors, poses a significant challenge to hazard prevention. Existing forecasting approache…
cs.LG2026
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando MartÃnez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
cs.LG2025
PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
Wanjia Zhao, Qinwei Ma, Jingzhe Shi +7
Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively explored. Most existing physics benchmarks eva…