15 papers
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +10
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs h…
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
Daksh Mittal, Tommaso Castellani, Thomson Yen +7
We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared acro…
Learning-To-Measure: In-Context Active Feature Acquisition
Yuta Kobayashi, Zilin Jing, Jiayu Yao +2
Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances by adaptively selecting which features to…
A Broader View of Thompson Sampling
Yanlin Qu, Hongseok Namkoong, Assaf Zeevi
Thompson Sampling is one of the most widely used and studied bandit algorithms, known for its simple structure, low regret performance, and solid theoretical guarantees. Yet, in st…
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
Tommaso Castellani, Naimeng Ye, Daksh Mittal +4
For agentic systems to use external tools to solve complex, long-horizon tasks, we need a large set of diverse and controllable tool-use environments. We introduce SynthTools, a fu…
A Sensitivity Approach to Causal Inference Under Limited Overlap
Yuanzhe Ma, Yian Huang, Hongseok Namkoong
Limited overlap between treated and control groups is a key challenge in observational analysis. Standard approaches like trimming importance weights can reduce variance but introd…