activity
20182026
collaborators

29 papers

cs.AI2026

JarvisBench: Always-on Intelligence Between Humans and Agents

Chen Chen, Zhehuai Chen

Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate acces…

eess.AS2026

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor +17

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as us…

cs.CL2026

Voice Memory for Agentic Speech Recognition

Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance wh…

cs.AI2026

Just A Rather Very Intelligent Spoken Agent

Chen Chen, Zhehuai Chen

Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, users give an initial instruction, receive…

eess.AS2026

Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning

Yuxin Li, Donghang Wu, Guan-Ting Lin +4

Recent full-duplex spoken dialogue models have demonstrated compelling progress toward human-like interaction, enabling agents to respond with low latency, produce backchannels, an…

eess.AS2026

Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

Guan-Ting Lin, Chen Chen, Zhehuai Chen +1

We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our d…