27 citations · 27 across the 1 of their papers we have counts for
7 papers
The Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap
Shota Horiguchi, Nelson Yalta, Paola Garcia +7
This paper provides a detailed description of the Hitachi-JHU system that was submitted to the Third DIHARD Speech Diarization Challenge. The system outputs the ensemble results of…
ESPnet-ST: All-in-One Speech Translation Toolkit
Hirofumi Inaguma, Shun Kiyono, Kevin Duh +4
We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech…
HATSUKI : An anime character like robot figure platform with anime-style expressions and imitation learning based action generation
Pin-Chu Yang, Mohammed Al-Sada, Chang-Chieh Chiu +6
Japanese character figurines are popular and have pivot position in Otaku culture. Although numerous robots have been developed, less have focused on otaku-culture or on embodying…
CNN-based MultiChannel End-to-End Speech Recognition for everyday home environments
Nelson Yalta, Shinji Watanabe, Takaaki Hori +2
Casual conversations involving multiple speakers and noises from surrounding devices are common in everyday environments, which degrades the performances of automatic speech recogn…
Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling
Jaejin Cho, Murali Karthick Baskar, Ruizhi Li +6
Sequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech research. The approach benefits by performing model training without using lexi…
Weakly Supervised Deep Recurrent Neural Networks for Basic Dance Step Generation
Nelson Yalta, Shinji Watanabe, Kazuhiro Nakadai +1
Synthesizing human's movements such as dancing is a flourishing research field which has several applications in computer graphics. Recent studies have demonstrated the advantages…