activity
20172024
most citedAuto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

128 citations · 128 across the 5 of their papers we have counts for

collaborators

7 papers

cs.SD2024

Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models

Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin +2

This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viabi…

cs.CL2024

Dynamic Data Pruning for Automatic Speech Recognition

Qiao Xiao, Pingchuan Ma, Adriana Fernandez-Lopez +7

The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitivel…

cs.CV2024

MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization

Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma +5

Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facil…

cs.CV2023

SparseVSR: Lightweight and Noise Robust Visual Speech Recognition

Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma +3

Recent advances in deep neural networks have achieved unprecedented success in visual speech recognition. However, there remains substantial disparity between current methods and t…

cs.CV2023★ 128 cited

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez +3

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speec…

cs.CV2017

Automatic Viseme Vocabulary Construction to Enhance Continuous Lip-reading

Adriana Fernandez-Lopez, Federico M. Sukno

Speech is the most common communication method between humans and involves the perception of both auditory and visual channels. Automatic speech recognition focuses on interpreting…