4 papers
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
Tuyen Tran, Thao Minh Le, Quang-Hung Le +1
Vision-language alignment in video must address the complexity of language, evolving interacting entities, their action chains, and semantic gaps between language and vision. This…
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
Tuyen Tran, Thao Minh Le, Truyen Tran
Referring-based Video Object Segmentation is a multimodal problem that requires producing fine-grained segmentation results guided by external cues. Traditional approaches to this…
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
Khai Le-Duc, Tuyen Tran, Bach Phan Tat +10
Multilingual speech translation (ST) and machine translation (MT) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviat…
Unified Framework with Consistency across Modalities for Human Activity Recognition
Tuyen Tran, Thao Minh Le, Hung Tran +1
Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input m…