7 papers
Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions
Kun Zhou, Jiakai He, Wenmian Yang +3
Presentation slides are a primary medium for data-driven reporting, yet keeping complex, analytics-style decks up to date remains labor-intensive. Existing automation methods mostl…
ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification
Zhensheng Wang, ZhanTeng Lin, Wenmian Yang +3
The advancement of large language models (LLMs) has enhanced tabular question answering (Tabular QA), yet they struggle with open-domain queries exhibiting underspecified or uncert…
CocoaBench: Evaluating Unified Digital Agents in the Wild
CocoaBench Team, Shibo Hao, Zhining Zhang +29
LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…
The PokeAgent Challenge: Competitive and Long-Context Learning at Scale
Seth Karten, Jake Grigsby, Tersoo Upaa +28
We present the PokeAgent Challenge, a large-scale benchmark for decision-making research built on Pokemon's multi-agent battle system and expansive role-playing game (RPG) environm…
Codified Finite-state Machines for Role-playing
Letian Peng, Yupeng Hou, Kun Zhou +1
Modeling latent character states is crucial for consistent and engaging role-playing (RP) with large language models (LLMs). Yet, existing prompting-based approaches mainly capture…
Deriving Character Logic from Storyline as Codified Decision Trees
Letian Peng, Kun Zhou, Longfei Yun +2
Role-playing (RP) agents rely on behavioral profiles to act consistently across diverse narrative contexts, yet existing profiles are largely unstructured, non-executable, and weak…