4 papers
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
Yanda Li, Yuhan Liu, Zirui Song +3
Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be un…
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
Chong Tian, Yu Wang, Chenxu Yang +5
Short-form video platforms are major channels for news but also fertile ground for multimodal misinformation where each modality appears plausible alone yet cross-modal relationshi…
MMoE: Robust Spoiler Detection with Multi-modal Information and Domain-aware Mixture-of-Experts
Zinan Zeng, Sen Ye, Zijian Cai +4
Online movie review websites are valuable for information and discussion about movies. However, the massive spoiler reviews detract from the movie-watching experience, making spoil…
Rate-Aware Learned Speech Compression
Jun Xu, Zhengxue Cheng, Guangchuan Chi +3
The rapid rise of real-time communication and large language models has significantly increased the importance of speech compression. Deep learning-based neural speech codecs have…