9 citations · 26 across the 12 of their papers we have counts for
17 papers
Tagged End-to-End Simultaneous Speech Translation Training using Simultaneous Interpretation Data
Yuka Ko, Ryo Fukuda, Yuta Nishikawa +3
Simultaneous speech translation (SimulST) translates partial speech inputs incrementally. Although the monotonic correspondence between input and output is preferable for smaller l…
E2E Refined Dataset
Keisuke Toyama, Katsuhito Sudoh, Satoshi Nakamura
Although the well-known MR-to-text E2E dataset has been used by many researchers, its MR-text pairs include many deletion/insertion/substitution errors. Since such errors affect th…
vTTS: visual-text to speech
Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi +2
This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into…
Simultaneous Neural Machine Translation with Constituent Label Prediction
Yasumasa Kano, Katsuhito Sudoh, Satoshi Nakamura
Simultaneous translation is a task in which translation begins before the speaker has finished speaking, so it is important to decide when to start the translation process. However…
Using Perturbed Length-aware Positional Encoding for Non-autoregressive Neural Machine Translation
Yui Oka, Katsuhito Sudoh, Satoshi Nakamura
Non-autoregressive neural machine translation (NAT) usually employs sequence-level knowledge distillation using autoregressive neural machine translation (AT) as its teacher model.…
ARTA: Collection and Classification of Ambiguous Requests and Thoughtful Actions
Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh +1
Human-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, e…