collaborators

6 papers

cs.AI2026

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

Mina Remeli, Moritz Hardt

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues…

cs.LG2026

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Shashwat Goel, Nikhil Chandak, Arvindh Arun +5

AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for rea…

cs.AI2026

Computational Arbitrage in AI Model Markets

Ricardo Olmedo, Bernhard Schölkopf, Moritz Hardt

Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a…

cs.LG2026

Scaling Open-Ended Reasoning to Predict the Future

Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2

High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. T…

cs.LG2025

Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning

Jonas Hübotter, Leander Diaz-Bone, Ido Hakimi +2

Humans are good at learning on the job: We learn how to solve the tasks we face as we go along. Can a model do the same? We propose an agent that assembles a task-specific curricul…

cs.CL2025

Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2

Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple ch…