23 papers
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Kaiwen Xiong, Haonian Ji, Shi Qiu +4
Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and…
VisualClaw: A Real-Time, Personalized Agent for the Physical World
Haoqin Tu, Jianwen Chen, Zijun Wang +14
Vision language models are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cos…
Not All Skills Help: Measuring and Repairing Agent Knowledge
Yixuan Wang, Yiyang Zhou, Yiming Liang +4
LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every decision about which skills to keep and how…
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
Yuxiang Lai, Peng Xia, Haonian Ji +8
Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static pr…
ClawArena: Benchmarking AI Agents in Evolving Information Environments
Haonian Ji, Kaiwen Xiong, Siwei Han +9
AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is scattered across heterogeneous sources…