activity
20242026
collaborators

13 papers

cs.AI2026

How to Interpret Agent Behavior

Jie Gao, Kaiser Sun, Jen-tse Huang +8

Autonomous agents such as Claude Code and Codex now operate for hours or even days. Understanding their runtime behavior has become critical for downstream tasks such as diagnosing…

cs.CV2025

Generative World Explorer

Taiming Lu, Tianmin Shu, Alan Yuille +2

Planning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their en…

cs.CL2025

WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment

Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme +1

AI systems make decisions in physical environments through primitive actions or affordances that are accessed via API calls. While deploying AI agents in the real world involves nu…

cs.CL2025

Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data

Jingyu Zhang, Marc Marone, Tianjian Li +2

To trust the fluent generations of large language models (LLMs), humans must be able to verify their correctness against trusted, external sources. Recent efforts, such as providin…

cs.AI2025

Tur[k]ingBench: A Challenge Benchmark for Web Agents

Kevin Xu, Yeganeh Kordi, Tanay Nayak +7

Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-ta…

cs.CL2025

Benchmarking Language Model Creativity: A Case Study on Code Generation

Yining Lu, Dixuan Wang, Tianjian Li +4

As LLMs become increasingly prevalent, it is interesting to consider how ``creative'' these models can be. From cognitive science, creativity consists of at least two key character…