3 citations · 5 across the 4 of their papers we have counts for
4 papers
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
Yufeng Yang, Desh Raj, Ju Lin +8
The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. Howeve…
End-to-End Speech Recognition Contextualization with Large Language Models
Egor Lakomkin, Chunyang Wu, Yassir Fathullah +3
In recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities.…
Prompting Large Language Models with Speech Recognition Abilities
Yassir Fathullah, Chunyang Wu, Egor Lakomkin +9
Large language models have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. I…
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision
Xubo Liu, Egor Lakomkin, Konstantinos Vougioukas +9
Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video…