4 papers · 1 filter
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Kaiwen Xiong, Haonian Ji, Shi Qiu +4
Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and…
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Jiaqi Liu, Shi Qiu, Mairui Li +33
Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail…
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
Yuxiang Lai, Peng Xia, Haonian Ji +8
Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static pr…
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
Haonian Ji, Shi Qiu, Siyang Xin +5
While foundation models (FMs), such as diffusion models and large vision-language models (LVLMs), have been widely applied in educational contexts, their ability to generate pedago…