10 papers
An Algebraic Method for Optimizing the State, Control, and Terminal State Weight Matrices for Optimal Feedback Control
Daegyun Choi, Donghoon Kim, James D. Turner
The necessary conditions for formulating optimal feedback control algorithms have been known for many years. Free parameters exist in the performance index in the form of state and…
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
Ahmad Salimi, Wentao Ma, Yuzhi Tang +3
Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress throu…
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Yuzhi Tang, Wentao Ma, Xiling Zhao +17
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed…
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
Rasool Fakoor, Murdock Aubry, Nicholas Stranges +1
Reinforcement learning is structurally harder than supervised learning because the policy changes the data distribution it learns from. The resulting fragility is especially visibl…
ProactBench: Beyond What The User Asked For
Sepehr Harfi, Ahmad Salimi, Dongming Shen +1
Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implie…
The Pokémon Theorem and other Fairness Impossibility Results
Daniel Matsui Smola, Alex Smola
Fairness impossibility results often look like distinct scalar incompatibility statements. We show that several share one RKHS geometry: fairness criteria are linear constraints on…