Showing eess.ASShow all
2 papers · 1 filter
eess.AS2026
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
Haolin He, Xingjian Du, Renhe Sun +16
Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing atten…
eess.AS2024
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
Dehua Tao, Daxin Tan, Yu Ting Yeung +2
Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech syn…