4 papers
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Yuzhi Tang, Wentao Ma, Xiling Zhao +17
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed…
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
Ahmad Salimi, Wentao Ma, Yuzhi Tang +3
Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress throu…
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
Pengfei Yu, Dongming Shen, Silin Meng +8
We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…
Optimal Control of Partially Observable Markov Decision Processes with Finite Linear Temporal Logic Constraints
Krishna C. Kalagarla, Dhruva Kartik, Dongming Shen +3
Autonomous agents often operate in scenarios where the state is partially observed. In addition to maximizing their cumulative reward, agents must execute complex tasks with rich t…