From the 1 of 14 linked papers with an AI index.
14 papers
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning
Tao Zhou, Jing Han, Lingyu Shu +1
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing approaches struggle to capture the inherentl…
CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection
Qiyang Sun, Yi Chang, Yupei Li +3
The paper introduces CHARM, a training-free framework that calibrates large language model predictions and fuses prosodic acoustic cues to improve zero-shot multimodal sarcasm dete…
SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition
Qiyang Sun, Yi Chang, Zixing Zhang +1
Speech conveys rich emotional information. As Speech Emotion Recognition (SER) is usually deployed in privacy-sensitive and reliability-critical environments, adversarial attacks o…
AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
Cheng Zhu, Jing Han, Qianshuai Xue +3
Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient f…
Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis
Zhiqiang Gao, Shihao Gao, Zixing Zhang +3
Understanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-base…
PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
Bin Wang, Yang Xu, Huan Zhao +2
Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high…