3 papers
eess.AS2026
Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity
Zhenwei Mou, Liping Chen, Yajun Hu +3
Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based…
eess.AS2024
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
Minghui Wu, Luzhen Xu, Jie Zhang +15
This report describes the submitted system to the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) challenge, which considers the ASR task with multi-speaker overlappin…
cs.SD2024
Multitask frame-level learning for few-shot sound event detection
Liang Zou, Genwei Yan, Ruoyu Wang +4
This paper focuses on few-shot Sound Event Detection (SED), which aims to automatically recognize and classify sound events with limited samples. However, prevailing methods method…