activity
20162022
most citedCHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

97 citations · 580 across the 99 of their papers we have counts for

collaborators

152 papers

eess.AS20221 cited

SpeechLMScore: Evaluating speech generation using speech language model

Soumi Maiti, Yifan Peng, Takaaki Saeki +1

While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality…

eess.AS2022

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody…

cs.CL2022

Align, Write, Re-order: Explainable End-to-End Speech Translation via Operation Sequence Generation

Motoi Omachi, Brian Yan, Siddharth Dalmia +2

The black-box nature of end-to-end speech translation (E2E ST) systems makes it difficult to understand how source language inputs are being mapped to the target language. To solve…

cs.CL2022

A Study on the Integration of Pre-trained SSL, ASR, LM and SLU Models for Spoken Language Understanding

Yifan Peng, Siddhant Arora, Yosuke Higuchi +6

Collecting sufficient labeled data for spoken language understanding (SLU) is expensive and time-consuming. Recent studies achieved promising results by using pre-trained models in…

cs.CL2022

Towards Zero-Shot Code-Switched Speech Recognition

Brian Yan, Matthew Wiesner, Ondrej Klejch +2

In this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot setting where no transcribed CS speech data is available…

cs.CL20221 cited

Bridging Speech and Textual Pre-trained Models with Unsupervised ASR

Jiatong Shi, Chan-Jan Hsu, Holam Chung +5

Spoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised mode…