31 citations · 32 across the 5 of their papers we have counts for
5 papers
LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
Junki Ohmura, Yuki Ito, Emiru Tsunoo +2
Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impr…
Polyphone disambiguation and accent prediction using pre-trained language models in Japanese TTS front-end
Rem Hida, Masaki Hamada, Chie Kamada +3
Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information fr…
Towards Online End-to-end Transformer Automatic Speech Recognition
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura +1
The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR…
Transformer ASR with Contextual Block Processing
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura +1
The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks (RNNs) in end-to-end (E2E) automatic speech recogniti…
End-to-end Adaptation with Backpropagation through WFST for On-device Speech Recognition System
Emiru Tsunoo, Yosuke Kashiwagi, Satoshi Asakawa +1
An on-device DNN-HMM speech recognition system efficiently works with a limited vocabulary in the presence of a variety of predictable noise. In such a case, vocabulary and environ…