13 papers
MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis
Yuhua Wen, Yingying Zhou, Qifei Li +4
Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted mo…
CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang +7
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoi…
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
Yizhong Geng, Wenxin Fu, Kecan Mao +7
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integratio…
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
Zhaoyang Meng, Zhengyao Ma, Kecan Mao +2
Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, whic…
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
Cong Wang, Yizhong Geng, Yuhua Wen +7
Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce…
Hello-Chat: Towards Realistic Social Audio Interactions
Yueran Hou, Peilei Jia, Zihan Sun +5
Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer fr…