3 papers
cs.CV2026
Head-Pose-Aware Visual Speech Recognition with FiLM Modulation
Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh
Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme ambiguity and pose-induced v…
cs.CV2026
Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction
Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh
Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expres…
cs.CV2025
Leveraging Generalizability of Image-to-Image Translation for Enhanced Adversarial Defense
Haibo Zhang, Zhihua Yao, Kouichi Sakurai +1
In the rapidly evolving field of artificial intelligence, machine learning emerges as a key technology characterized by its vast potential and inherent risks. The stability and rel…