6 papers
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
Junyoung Seo, Rodrigo Mira, Alexandros Haliassos +6
Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One soluti…
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
Umberto Cappellazzo, Minsu Kim, Pingchuan Ma +4
Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granulari…
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
Minsu Kim, Pingchuan Ma, Honglie Chen +2
This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace,…
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
Minsu Kim, Rodrigo Mira, Honglie Chen +2
In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Con…
Large Language Models are Strong Audio-Visual Speech Recognition Learners
Umberto Cappellazzo, Minsu Kim, Honglie Chen +5
Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and…
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
Alexandros Haliassos, Rodrigo Mira, Honglie Chen +3
Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised stud…