activity
20192021
most citedNeural Speaker Diarization with Speaker-Wise Chain Rule

41 citations · 86 across the 13 of their papers we have counts for

collaborators

19 papers

cs.CL2021

Team Hitachi @ AutoMin 2021: Reference-free Automatic Minuting Pipeline with Argument Structure Construction over Topic-based Summarization

Atsuki Yamaguchi, Gaku Morio, Hiroaki Ozaki +2

This paper introduces the proposed automatic minuting system of the Hitachi team for the First Shared Task on Automatic Minuting (AutoMin-2021). We utilize a reference-free approac…

cs.RO2021

Emotional Speech Synthesis for Companion Robot to Imitate Professional Caregiver Speech

Takeshi Homma, Qinghua Sun, Takuya Fujioka +5

When people try to influence others to do something, they subconsciously adjust their speech to include appropriate emotional information. In order for a robot to influence people…

eess.AS2021

Semi-Supervised Training with Pseudo-Labeling for End-to-End Neural Diarization

Yuki Takashima, Yusuke Fujita, Shota Horiguchi +3

In this paper, we present a semi-supervised training technique using pseudo-labeling for end-to-end neural diarization (EEND). The EEND system has shown promising performance compa…

eess.AS2021

End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection

Yuki Takashima, Yusuke Fujita, Shinji Watanabe +3

In this paper, we present a conditional multitask learning method for end-to-end neural speaker diarization (EEND). The EEND system has shown promising performance compared with tr…

cs.SD2021★ 2 cited

Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers

Yawen Xue, Shota Horiguchi, Yusuke Fujita +4

We propose a streaming diarization method based on an end-to-end neural diarization (EEND) model, which handles flexible numbers of speakers and overlapping speech. In our previous…

eess.AS2020★ 3 cited

Building Multi lingual TTS using Cross Lingual Voice Conversion

Qinghua Sun, Kenji Nagamatsu

In this paper we propose a new cross-lingual Voice Conversion (VC) approach which can generate all speech parameters (MCEP, LF0, BAP) from one DNN model using PPGs (Phonetic Poster…