3 citations · 3 across the 4 of their papers we have counts for
4 papers
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Jinzheng Zhao, Niko Moritz, Egor Lakomkin +7
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumu…
Effective internal language model training and fusion for factorized transducer model
Jinxi Guo, Niko Moritz, Yingyi Ma +6
The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracte…
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
Ju Lin, Niko Moritz, Yiteng Huang +4
Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introdu…
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision
Xubo Liu, Egor Lakomkin, Konstantinos Vougioukas +9
Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video…