activity
20242026
collaborators

14 papers

cs.LG2026

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Shashwat Goel, Nikhil Chandak, Arvindh Arun +5

AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for rea…

cs.AI2026

The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs

Akshit Sinha, Arvindh Arun, Shashwat Goel +2

Does continued scaling of large language models (LLMs) yield diminishing returns? In this work, we show that short-task benchmarks may give an illusion of slowing progress, as even…

cs.LG2026

Intrinsic Credit Assignment for Long Horizon Interaction

Ilze Amanda Auzina, Joschka Strüber, Sergio Hernández-Gutiérrez +3

How can we train agents to navigate uncertainty over long horizons? In this work, we propose ΔBelief-RL, which leverages a language model's own intrinsic beliefs to reward interme…

cs.LG2026

Scaling Open-Ended Reasoning to Predict the Future

Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2

High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. T…

cs.LG2025

Training AI Co-Scientists Using Rubric Rewards

Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8

AI co-scientists are emerging as a tool to assist human researchers in achieving their research goals. A crucial feature of these AI co-scientists is the ability to generate a rese…

cs.CL2025

What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages

Debangan Mishra, Arihant Rastogi, Agyeya Negi +2

How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric applied to 20 languages and 47 subject…