1 citations · 1 across the 5 of their papers we have counts for
5 papers · 1 filter
EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
Haiyang Shen, Xuanzhong Chen, Wendong Xu +3
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This leaves out a basi…
MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis
Haiyang Shen, Taian Guo, Xuanzhong Chen +11
Although LLMs have made substantial progress in reasoning, systematically producing frontier-level reasoning data remains difficult. Existing synthesis methods often have limited v…
MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
Guoxin Chen, Zile Qiao, Wenqing Wang +10
Large Reasoning Models (LRMs) face two fundamental limitations: excessive token consumption when overanalyzing simple information processing tasks, and inability to access up-to-da…
IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
Guoxin Chen, Zile Qiao, Xuanzhong Chen +13
Recent advances in deep-research agents have shown promise for autonomous knowledge construction through dynamic reasoning over external sources. However, existing approaches rely…
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Rui Min, Zile Qiao, Ze Xu +18
Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. Whi…