works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.MM2026

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4

The paper introduces Flowley, an end‑to‑end model that generates synchronized audio directly from silent video using a novel progressive soft‑masked cross‑attention mechanism, and…

cs.CV2026

Vortex: Multi-Modal Fusion System for Intelligent Video Retrieval

Duc-Tho Nguyen, Hieu-Hoc Tran-Minh, Khanh-Hoa Lam +4

This paper presents Vortex, the multimodal video retrieval system developed by our team, FocusOnFun, for the Ho Chi Minh City AI Challenge 2025, designed to advance intelligent mul…

cs.SD2026

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen +2

Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregres…

cs.CV2026

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi +3

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To addres…

cs.SD2025

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham +1

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content,…

cs.CL2024

Effective Context Modeling Framework for Emotion Recognition in Conversations

Cuong Tran Van, Thanh V. T. Tran, Van Nguyen +1

Emotion Recognition in Conversations (ERC) facilitates a deeper understanding of the emotions conveyed by speakers in each utterance within a conversation. Recently, Graph Neural N…