3 papers
cs.SD2026
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
Yanda Li, Yuhan Liu, Zirui Song +3
Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be un…
cs.AI2026
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
Chong Tian, Yu Wang, Chenxu Yang +5
Short-form video platforms are major channels for news but also fertile ground for multimodal misinformation where each modality appears plausible alone yet cross-modal relationshi…
eess.AS2025
Rate-Aware Learned Speech Compression
Jun Xu, Zhengxue Cheng, Guangchuan Chi +3
The rapid rise of real-time communication and large language models has significantly increased the importance of speech compression. Deep learning-based neural speech codecs have…