5 papers
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +10
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs h…
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
Daksh Mittal, Tommaso Castellani, Thomson Yen +7
We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared acro…
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
Tommaso Castellani, Naimeng Ye, Daksh Mittal +4
For agentic systems to use external tools to solve complex, long-horizon tasks, we need a large set of diverse and controllable tool-use environments. We introduce SynthTools, a fu…
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
Gilbert Yang, Yaqin Chen, Thomson Yen +1
To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience. However, current evaluations of LLM-…
Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
Thomson Yen, Andrew Wei Tung Siah, Haozhe Chen +3
Careful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, maki…