From the 1 of 7 linked papers with an AI index.
7 papers
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4
The paper introduces Flowley, an end‑to‑end model that generates synchronized audio directly from silent video using a novel progressive soft‑masked cross‑attention mechanism, and…
Vortex: Multi-Modal Fusion System for Intelligent Video Retrieval
Duc-Tho Nguyen, Hieu-Hoc Tran-Minh, Khanh-Hoa Lam +4
This paper presents Vortex, the multimodal video retrieval system developed by our team, FocusOnFun, for the Ho Chi Minh City AI Challenge 2025, designed to advance intelligent mul…
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen +2
Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregres…
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi +3
Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To addres…
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham +1
Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content,…
Effective Context Modeling Framework for Emotion Recognition in Conversations
Cuong Tran Van, Thanh V. T. Tran, Van Nguyen +1
Emotion Recognition in Conversations (ERC) facilitates a deeper understanding of the emotions conveyed by speakers in each utterance within a conversation. Recently, Graph Neural N…