3 papers
eess.AS2026
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
Haoyuan Yang, Mu Yang, Jiamin Xie +2
Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacit…
eess.AS2025
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
Mu Yang, Szu-Jui Chen, Jiamin Xie +1
One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete token-based para…
eess.AS2025
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang +8
Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech L…