collaborators

5 papers

cs.AI2026

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Xirui Li, Ming Li, Ion Stoica +2

Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that what is needed is not just a dat…

cs.AI2026

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

Zhonghao Yang, Yu Li, Yanxu Zhu +6

As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve with them. ATBench is a diverse…

cs.AI2026

Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs

Paiheng Xu, Gang Wu, Xiang Chen +6

Scripting interfaces enable users to automate tasks and customize software workflows, but creating scripts traditionally requires programming expertise and familiarity with specifi…

cs.CL2025

Multiple LLM Agents Debate for Equitable Cultural Alignment

Dayeon Ki, Rachel Rudinger, Tianyi Zhou +1

Large Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on si…

cs.AI2025

GraphicBench: A Planning Benchmark for Graphic Design with Language Agents

Dayeon Ki, Tianyi Zhou, Marine Carpuat +3

Large Language Model (LLM)-powered agents have unlocked new possibilities for automating human tasks. While prior work has focused on well-defined tasks with specified goals, the c…