7 papers
Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets
Xuan Liu, Haoyang Shang, Haojian Jin
AI agents are emerging as market participants that trade delegated cognitive work with one another on behalf of their users. Each agent can act as both a task poster and a contract…
Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents
Xuan Liu, HaoYang Shang, Zizhang Liu +5
We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…
Understanding Parents' Desires in Moderating Children's Interactions with GenAI Chatbots through LLM-Generated Probes
John Driscoll, Yulin Chen, Viki Shi +3
This paper studies how parents want to moderate children's interactions with Generative AI chatbots, with the goal of informing the design of future GenAI parental control tools. W…
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments
Yuxuan Zhang, Haoyang Yu, Lanxiang Hu +2
We introduce a modular harness design for LLM agents that composes of perception, memory, and reasoning components, enabling a single LLM or VLM backbone to tackle a wide spectrum…
lmgame-Bench: How Good are LLMs at Playing Games?
Lanxiang Hu, Mingjia Huo, Yuxuan Zhang +6
Playing video games requires perception, memory, and planning, exactly the faculties modern large language model (LLM) agents are expected to master. We study the major challenges…
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
Yanbin Yin, Kun Zhou, Zhen Wang +11
The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard prac…