activity
20172026
most citedTraining Strategies for Improved Lip-reading

59 citations · 93 across the 34 of their papers we have counts for

collaborators
Showing cs.CVShow all

33 papers · 1 filter

cs.CV2026

Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition

Alexandros Haliassos, Rodrigo Mira, Stavros Petridis

Unified Speech Recognition (USR) has emerged as a semi-supervised framework for training a single model for audio, visual, and audiovisual speech recognition, achieving state-of-th…

cs.CV2025

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

Andreas Zinonos, Michał Stypułkowski, Antoni Bigata +3

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 1…

cs.CV2025

FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion

Kazuaki Mishima, Antoni Bigata Casademunt, Stavros Petridis +2

Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While rece…

cs.CV2025

KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution

Antoni Bigata, Rodrigo Mira, Stella Bounareli +4

Lip synchronization, known as the task of aligning lip movements in an existing video with new input audio, is typically framed as a simpler variant of audio-driven facial animatio…

cs.CV2025

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

Antoni Bigata, Michał Stypułkowski, Rodrigo Mira +7

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. E…

cs.CV2025

Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Umberto Cappellazzo, Minsu Kim, Stavros Petridis

Audio-Visual Speech Recognition (AVSR) leverages audio and visual modalities to improve robustness in noisy environments. Recent advances in Large Language Models (LLMs) show stron…