activity
20172022
most citedMultichannel End-to-end Speech Recognition

46 citations · 110 across the 14 of their papers we have counts for

collaborators

32 papers

cs.SD2022

Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR

Xuankai Chang, Niko Moritz, Takaaki Hori +2

Graph-based temporal classification (GTC), a generalized form of the connectionist temporal classification loss, was recently proposed to improve automatic speech recognition (ASR)…

cs.CL2021

Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning

Ankit P. Shah, Shijie Geng, Peng Gao +5

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…

eess.AS20211 cited

Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy

Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1

Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to e…

cs.CV2021

Optimizing Latency for Online Video CaptioningUsing Audio-Visual Transformers

Chiori Hori, Takaaki Hori, Jonathan Le Roux

Video captioning is an essential technology to understand scenes and describe events in natural language. To apply it to real-time monitoring, a system needs not only to describe e…

eess.AS2021

Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition

Niko Moritz, Takaaki Hori, Jonathan Le Roux

Attention-based end-to-end automatic speech recognition (ASR) systems have recently demonstrated state-of-the-art results for numerous tasks. However, the application of self-atten…

eess.AS20214 cited

Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition

Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1

Pseudo-labeling (PL) has been shown to be effective in semi-supervised automatic speech recognition (ASR), where a base model is self-trained with pseudo-labels generated from unla…