13 papers
How to Interpret Agent Behavior
Jie Gao, Kaiser Sun, Jen-tse Huang +8
Autonomous agents such as Claude Code and Codex now operate for hours or even days. Understanding their runtime behavior has become critical for downstream tasks such as diagnosing…
Generative World Explorer
Taiming Lu, Tianmin Shu, Alan Yuille +2
Planning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their en…
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme +1
AI systems make decisions in physical environments through primitive actions or affordances that are accessed via API calls. While deploying AI agents in the real world involves nu…
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
Jingyu Zhang, Marc Marone, Tianjian Li +2
To trust the fluent generations of large language models (LLMs), humans must be able to verify their correctness against trusted, external sources. Recent efforts, such as providin…
Tur[k]ingBench: A Challenge Benchmark for Web Agents
Kevin Xu, Yeganeh Kordi, Tanay Nayak +7
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-ta…
Benchmarking Language Model Creativity: A Case Study on Code Generation
Yining Lu, Dixuan Wang, Tianjian Li +4
As LLMs become increasingly prevalent, it is interesting to consider how ``creative'' these models can be. From cognitive science, creativity consists of at least two key character…