Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
James Mooney, Josef Woldense, Zheng Robert Jia +4
The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To…
cs.AI2026
The Amazing Agent Race: Strong Tool Users, Weak Navigators
Zae Myung Kim, Dongseok Lee, Jaehyung Kim +2
Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of 2 to 5 steps. We introduce…
cs.AI2025
BAID: A Benchmark for Bias Assessment of AI Detectors
Priyam Basu, Yunfeng Zhang, Vipul Raheja
AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against Englis…