Showing cs.MMShow all
3 papers · 1 filter
cs.MM2026
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4
Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress,…
cs.MM2025
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
Van-Hoang Phan, Tung-Duong Le-Duc, Long-Khanh Pham +7
The proliferation of multimedia content on social media platforms has dramatically transformed how information is consumed and disseminated. While this shift enables real-time cove…
cs.MM2025
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
Van-Hoang Phan, Long-Khanh Pham, Dang Vu +2
The rapid spread of misinformation in mobile and wireless networks presents critical security challenges. This study introduces a training-free, retrieval-based multimodal fact ver…