5 papers
Uncovering Insights of Compound Flooding with Data-Driven AI
Xu Zheng, Chaohao Lin, Sipeng Chen +7
Compound flooding, driven by nonlinear interactions between multiple hydrometeorological factors, poses a significant challenge to hazard prevention. Existing forecasting approache…
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando MartÃnez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
Wanjia Zhao, Qinwei Ma, Jingzhe Shi +7
Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively explored. Most existing physics benchmarks eva…
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
Stephen Fitz, Peter Romero, Steven Basart +2
Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consisten…
Towards Structurally Explainable Machine-Generated Text Detection: A Graph-Perspective Framework
Xu Zheng, Zhuomin Chen, Esteban Schafir +7
Despite the success of machine-generated text detectors, the black-box nature remains a critical limitation. Traditional explainability methods rely on token-level saliency, insuff…