3 papers
cs.AI2026
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Zhenzhen Ren, Jiyan He, Xinpeng Zhang +5
Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations t…
cs.AI2025
GTM: Simulating the World of Tools for AI Agents
Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian +4
The integration of external tools is pivotal for empowering Large Language Model (LLM) agents with real-world capabilities. However, training these agents through direct, continuou…
cs.CR2025
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
Zhenzhen Ren, GuoBiao Li, Sheng Li +2
Despite providing superior performance, open-source large language models (LLMs) are vulnerable to abusive usage. To address this issue, recent works propose LLM fingerprinting met…