15 citations · 27 across the 4 of their papers we have counts for
4 papers
Turning Up the Heat: Assessing 2-m Temperature Forecast Errors in AI Weather Prediction Models During Heat Waves
Kelsey E. Ennis, Elizabeth A. Barnes, Marybeth C. Arcodia +2
Extreme heat is the deadliest weather-related hazard in the United States. Furthermore, it is increasing in intensity, frequency, and duration, making skillful forecasts vital to p…
HCAST: Human-Calibrated Autonomy Software Tasks
David Rein, Joel Becker, Amy Deng +19
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du +10
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refe…
Carefully choose the baseline: Lessons learned from applying XAI attribution methods for regression tasks in geoscience
Antonios Mamalakis, Elizabeth A. Barnes, Imme Ebert-Uphoff
Methods of eXplainable Artificial Intelligence (XAI) are used in geoscientific applications to gain insights into the decision-making strategy of Neural Networks (NNs) highlighting…