most citedInterpretable Unified Language Checking

8 citations · 12 across the 9 of their papers we have counts for

collaborators

9 papers

cs.SD2023

QS-TTS: Towards Semi-Supervised Text-to-Speech Synthesis via Vector-Quantized Self-Supervised Speech Representation Learning

Haohan Guo, Fenglong Xie, Jiawen Kang +3

This paper proposes a novel semi-supervised TTS framework, QS-TTS, to improve TTS quality with lower supervised data requirements via Vector-Quantized Self-Supervised Speech Repres…

cs.SD2023

Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information

Jie Chen, Changhe Song, Deyi Tuo +4

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic in…

cs.SD20231 cited

MSStyleTTS: Multi-Scale Style Modeling with Hierarchical Context Information for Expressive Speech Synthesis

Shun Lei, Yixuan Zhou, Liyang Chen +4

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the sty…

cs.SD20231 cited

Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator

Lingwei Meng, Jiawen Kang, Mingyu Cui +3

Multi-talker overlapped speech poses a significant challenge for speech recognition and diarization. Recent research indicated that these two tasks are inter-dependent and compleme…

cs.CL20238 cited

Interpretable Unified Language Checking

Tianhua Zhang, Hongyin Luo, Yung-Sung Chuang +7

Despite recent concerns about undesirable behaviors generated by large language models (LLMs), including non-factual, biased, and hateful language, we find LLMs are inherent multi-…

eess.AS2023

A Hierarchical Regression Chain Framework for Affective Vocal Burst Recognition

Jinchao Li, Xixin Wu, Kaitao Song +3

As a common way of emotion signaling via non-linguistic vocalizations, vocal burst (VB) plays an important role in daily social interaction. Understanding and modeling human vocal…