5 papers
Trading Human Curation for Synthetic Augmentation in RLVR
Akshansh, Leonardo Rosa Rodrigues, Michael Korostelev +2
The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models. Each task requires a sandbox…
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
Kate M. Lubrano, Faisal Sayed, Ankita Rathod +4
Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly importa…
LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design
Marilyn Zhang, Tianfeng Chen, Fabián Barzuna +2
LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let them converge on good designs in…
A large-scale evaluation of commonsense knowledge in humans and large language models
Tuan Dung Nguyen, Duncan J. Watts, Mark E. Whiting
Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implici…
AI Hasn't Fixed Teamwork, But It Shifted Collaborative Culture: A Longitudinal Study in a Project-Based Software Development Organization (2023-2025)
Qing Xiao, Xinlan Emily Hu, Mark E. Whiting +3
When AI entered the workplace, many believed it could reshape teamwork as profoundly as it boosted individual productivity. Would AI finally ease the longstanding challenges of tea…