14 papers
FutureSim: Replaying World Events to Evaluate Adaptive Agents
Shashwat Goel, Nikhil Chandak, Arvindh Arun +5
AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for rea…
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
Akshit Sinha, Arvindh Arun, Shashwat Goel +2
Does continued scaling of large language models (LLMs) yield diminishing returns? In this work, we show that short-task benchmarks may give an illusion of slowing progress, as even…
Intrinsic Credit Assignment for Long Horizon Interaction
Ilze Amanda Auzina, Joschka Strüber, Sergio Hernández-Gutiérrez +3
How can we train agents to navigate uncertainty over long horizons? In this work, we propose ÎBelief-RL, which leverages a language model's own intrinsic beliefs to reward interme…
Scaling Open-Ended Reasoning to Predict the Future
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. T…
Training AI Co-Scientists Using Rubric Rewards
Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8
AI co-scientists are emerging as a tool to assist human researchers in achieving their research goals. A crucial feature of these AI co-scientists is the ability to generate a rese…
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
Debangan Mishra, Arihant Rastogi, Agyeya Negi +2
How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric applied to 20 languages and 47 subject…