5 papers
EpochX: Building the Infrastructure for an Emergent Agent Civilization
Huacan Wang, Chaofa Yuan, Xialie Zhuang +15
General-purpose technologies reshape economies less by improving individual tools than by enabling new ways to organize production and coordination. We believe AI agents are approa…
EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
Shuo Zhang, Chaofa Yuan, Ryan Guo +11
While LLM-based agents have shown promise for deep research, most existing approaches rely on fixed workflows that struggle to adapt to real-world, open-ended queries. Recent work…
CloneMem: Benchmarking Long-Term Memory for AI Clones
Sen Hu, Zhiyu Zhang, Yuxiang Wei +4
AI Clones aim to simulate an individual's thoughts and behaviors to enable long-term, personalized interaction, placing stringent demands on memory systems to model experiences, em…
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
Yifu Guo, Zishan Xu, Zhiyuan Yao +6
Existing multimodal reasoning models and frameworks suffer from fundamental architectural limitations: most lack the human-like ability to autonomously explore diverse reasoning pa…
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
Ziyi Ni, Huacan Wang, Shuo Zhang +15
Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evalua…