2 papers
cs.SD2026
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
Keita Goto, Takashi Maekaku, Jin Sakuma +3
Dual-mode self-supervised speech models (S3Ms), which jointly pre-trained in the offline and online mode, suffer from attention mismatch in streaming scenarios due to missing futur…
cs.MM2025
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
Yuyue Wang, Xin Cheng, Yihan Wu +3
The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being co…