12 papers
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
Ke Liu, Jiwei Wei, Wenyu Zhang +5
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely…
RAVE: Re-Allocating Visual Attention in Large Multimodal Models
Xi Leng, Xinhong Ma, Ziqiang Dong +4
Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-moda…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search
Paul Greyson, Zhichao Geng, Wei Zhang +1
Music search at the scale of Amazon Music presents a unique challenge: queries frequently deviate from indexed metadata due to misspellings, transpositions, and phonetic variations…
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
Ran Ran, Jiwei Wei, Shuchang Zhou +5
Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query…
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
Ke Liu, Jiwei Wei, Shuchang Zhou +5
Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery p…