Interpreting the Role of Visemes in Audio-Visual Speech Recognition
arXiv:2509.16023 · doi:10.1109/ASRU65441.2025.11434749
Abstract
Audio-Visual Speech Recognition (AVSR) models have surpassed their audio-only counterparts in terms of performance. However, the interpretability of AVSR systems, particularly the role of the visual modality, remains under-explored. In this paper, we apply several interpretability techniques to examine how visemes are encoded in AV-HuBERT a state-of-the-art AVSR model. First, we use t-distributed Stochastic Neighbour Embedding (t-SNE) to visualize learned features, revealing natural clustering driven by visual cues, which is further refined by the presence of audio. Then, we employ probing to show how audio contributes to refining feature representations, particularly for visemes that are visually ambiguous or under-represented. Our findings shed light on the interplay between modalities in AVSR and could point to new strategies for leveraging visual information to improve AVSR performance.
Accepted into Automatic Speech Recognition and Understanding- ASRU 2025
References in corpus (4)
- Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
- VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
- Probing phoneme, language and speaker information in unsupervised speech representations
- Uncovering the Visual Contribution in Audio-Visual Speech Recognition