23 citations · 235 across the 69 of their papers we have counts for
22 papers · 1 filter
SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu +1
Large language models (LLMs) have gained considerable attention for Artificial Intelligence Generated Content (AIGC), particularly with the emergence of ChatGPT. However, the direc…
MiniSUPERB: Lightweight Benchmark for Self-supervised Speech Models
Yu-Hsiang Wang, Huang-Yu Chen, Kai-Wei Chang +2
SUPERB was proposed to evaluate the generalizability of self-supervised learning (SSL) speech models across various tasks. However, it incurs high computational costs due to the la…
Partially Fake Audio Detection by Self-attention-based Fake Span Discovery
Haibin Wu, Heng-Cheng Kuo, Naijun Zheng +5
The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly…
Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning
Yi-Chen Chen, Shu-wen Yang, Cheng-Kuang Lee +2
Speech representation learning plays a vital role in speech processing. Among them, self-supervised learning (SSL) has become an important research direction. It has been shown tha…
Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
Chung-Ming Chien, Jheng-Hao Lin, Chien-yu Huang +2
The few-shot multi-speaker multi-style voice cloning task is to synthesize utterances with voice and speaking style similar to a reference speaker given only a few reference sample…
Utilizing Self-supervised Representations for MOS Prediction
Wei-Cheng Tseng, Chien-yu Huang, Wei-Tsung Kao +2
Speech quality assessment has been a critical issue in speech processing for decades. Existing automatic evaluations usually require clean references or parallel ground truth data,…