5 papers
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama
Tianyang Xu, Hongqiu Wu, Weiqi Wu +1
LLM-based Interactive Drama introduces a novel dialogue scenario in which the player immerses into a character and engages in a dramatic story by interacting with LLM agents. Despi…
Towards Enhanced Immersion and Agency for LLM-based Interactive Drama
Hongqiu Wu, Weiqi Wu, Tianyang Xu +2
LLM-based Interactive Drama is a novel AI-based dialogue scenario, where the user (i.e. the player) plays the role of a character in the story, has conversations with characters pl…
Open Role-Playing with Delta-Engines
Hongqiu Wu, Zekai Xu, Tianyang Xu +5
Game roles can be reflections of personas from a parallel world. In this paper, we propose a new style of game-play to bridge self-expression and role-playing: \emph{open role-play…