4 citations · 10 across the 9 of their papers we have counts for
7 papers · 1 filter
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Jinzheng Zhao, Niko Moritz, Egor Lakomkin +7
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumu…
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
Yufeng Yang, Desh Raj, Ju Lin +8
The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. Howeve…
Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1
Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to e…
Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition
Niko Moritz, Takaaki Hori, Jonathan Le Roux
Attention-based end-to-end automatic speech recognition (ASR) systems have recently demonstrated state-of-the-art results for numerous tasks. However, the application of self-atten…
Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1
Pseudo-labeling (PL) has been shown to be effective in semi-supervised automatic speech recognition (ASR), where a base model is self-trained with pseudo-labels generated from unla…
Capturing Multi-Resolution Context by Dilated Self-Attention
Niko Moritz, Takaaki Hori, Jonathan Le Roux
Self-attention has become an important and widely used neural network component that helped to establish new state-of-the-art results for various applications, such as machine tran…