97 citations · 580 across the 99 of their papers we have counts for
152 papers
SpeechLMScore: Evaluating speech generation using speech language model
Soumi Maiti, Yifan Peng, Takaaki Saeki +1
While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality…
A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units
Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky
We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody…
Align, Write, Re-order: Explainable End-to-End Speech Translation via Operation Sequence Generation
Motoi Omachi, Brian Yan, Siddharth Dalmia +2
The black-box nature of end-to-end speech translation (E2E ST) systems makes it difficult to understand how source language inputs are being mapped to the target language. To solve…
A Study on the Integration of Pre-trained SSL, ASR, LM and SLU Models for Spoken Language Understanding
Yifan Peng, Siddhant Arora, Yosuke Higuchi +6
Collecting sufficient labeled data for spoken language understanding (SLU) is expensive and time-consuming. Recent studies achieved promising results by using pre-trained models in…
Towards Zero-Shot Code-Switched Speech Recognition
Brian Yan, Matthew Wiesner, Ondrej Klejch +2
In this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot setting where no transcribed CS speech data is available…
Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
Jiatong Shi, Chan-Jan Hsu, Holam Chung +5
Spoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised mode…