collaborators

5 papers

cs.CL2025

Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

Siddhant Arora, Haidar Khan, Kai Sun +14

End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with sig…

cs.LG2025

PIRF: Physics-Informed Reward Fine-Tuning for Diffusion Models

Mingze Yuan, Pengfei Jin, Na Li +1

Diffusion models have demonstrated strong generative capabilities across scientific domains, but often produce outputs that violate physical laws. We propose a new perspective by f…

cs.SD2025

MBCodec:Thorough disentangle for high-fidelity audio compression

Ruonan Zhang, Xiaoyang Hao, Yichen Han +3

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced…

eess.AS2025

Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses

Yufeng Yang, Yiteng Huang, Yong Xu +9

With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-co…

cs.RO2025

Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations

Haitong Ma, Bo Dai, Zhaolin Ren +2

Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducin…