17 citations · 29 across the 10 of their papers we have counts for
9 papers
Align, Write, Re-order: Explainable End-to-End Speech Translation via Operation Sequence Generation
Motoi Omachi, Brian Yan, Siddharth Dalmia +2
The black-box nature of end-to-end speech translation (E2E ST) systems makes it difficult to understand how source language inputs are being mapped to the target language. To solve…
End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation
Xuankai Chang, Takashi Maekaku, Yuya Fujita +1
This work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech…
A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation
Yosuke Higuchi, Nanxin Chen, Yuya Fujita +6
Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to aut…
Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
Tianzi Wang, Yuya Fujita, Xuankai Chang +1
Non-autoregressive (NAR) modeling has gained more and more attention in speech processing. With recent state-of-the-art attention-based automatic speech recognition (ASR) structure…
Toward Streaming ASR with Non-Autoregressive Insertion-based Model
Yuya Fujita, Tianzi Wang, Shinji Watanabe +1
Neural end-to-end (E2E) models have become a promising technique to realize practical automatic speech recognition (ASR) systems. When realizing such a system, one important issue…
Insertion-Based Modeling for End-to-End Automatic Speech Recognition
Yuya Fujita, Shinji Watanabe, Motoi Omachi +1
End-to-end (E2E) models have gained attention in the research field of automatic speech recognition (ASR). Many E2E models proposed so far assume left-to-right autoregressive gener…