3 papers
cs.RO2026
Training and Evaluating Diffusion Policies with Long Context Lengths
Abhinav Agarwal, Adam Wei, Taylan Kargin +6
Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, typically condition robot actions on only…
eess.AS2024
CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
Leying Zhang, Yao Qian, Long Zhou +9
Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along w…
cs.CL2024
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
Chenyang Le, Yao Qian, Dongmei Wang +8
There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most e…