From the 1 of 4 linked papers with an AI index.
4 papers
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Zhenzhen Ren, Jiyan He, Xinpeng Zhang +5
The paper introduces OrchBench, a deterministic simulation benchmark that evaluates multi‑agent orchestration plans on DAG‑structured tasks in isolation, providing fast, token‑effi…
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Haoxiang Guan, Jiyan He, Liyang Fan +10
Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (…
GTM: Simulating the World of Tools for AI Agents
Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian +4
The integration of external tools is pivotal for empowering Large Language Model (LLM) agents with real-world capabilities. However, training these agents through direct, continuou…
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
Zhenzhen Ren, GuoBiao Li, Sheng Li +2
Despite providing superior performance, open-source large language models (LLMs) are vulnerable to abusive usage. To address this issue, recent works propose LLM fingerprinting met…