From the 1 of 4 linked papers with an AI index.
4 papers
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4
The paper introduces Flowley, an end‑to‑end model that generates synchronized audio directly from silent video using a novel progressive soft‑masked cross‑attention mechanism, and…
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
Van-Hoang Phan, Tung-Duong Le-Duc, Long-Khanh Pham +7
The proliferation of multimedia content on social media platforms has dramatically transformed how information is consumed and disseminated. While this shift enables real-time cove…
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
Van-Hoang Phan, Long-Khanh Pham, Dang Vu +2
The rapid spread of misinformation in mobile and wireless networks presents critical security challenges. This study introduces a training-free, retrieval-based multimodal fact ver…
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham +1
Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content,…