2 citations · 4 across the 10 of their papers we have counts for
4 papers · 1 filter
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Sky Ng, Brihi Joshi, Ishan Gupta +47
Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain i…
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Yi Zhong, Buqiang Xu, Yijun Wang +4
At present, executable visual workflows have emerged as a mainstream paradigm in real-world industrial deployments, offering strong reliability and controllability. However, in cur…
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Jize Wang, Xuanxuan Liu, Yining Li +7
The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use be…
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Xiaomin Li, Mingye Gao, Yuexing Hao +4
Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making.…