From the 1 of 6 linked papers with an AI index.
6 papers
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Zhenzhen Ren, Jiyan He, Xinpeng Zhang +5
The paper introduces OrchBench, a deterministic simulation benchmark that evaluates multi‑agent orchestration plans on DAG‑structured tasks in isolation, providing fast, token‑effi…
Learning behavior accounts for background-related advantage in AI-assisted education
Jingwei Yi, Yueqi Xie, Jiyan He +7
Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average ef…
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Haoxiang Guan, Jiyan He, Liyang Fan +10
Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (…
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
Yanxi Wang, Zhiling Zhang, Wenbo Zhou +6
As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locat…
GTM: Simulating the World of Tools for AI Agents
Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian +4
The integration of external tools is pivotal for empowering Large Language Model (LLM) agents with real-world capabilities. However, training these agents through direct, continuou…
Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task Learning
Yuxuan Ren, Dihan Zheng, Chang Liu +7
In recent years, machine learning has demonstrated impressive capability in handling molecular science tasks. To support various molecular properties at scale, machine learning mod…