7 papers · 1 filter
When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery
Y Huynh, Duc Thanh Nguyen, Thao Minh Le +1
Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to comp…
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
Tuyen Tran, Thao Minh Le, Quang-Hung Le +1
Vision-language alignment in video must address the complexity of language, evolving interacting entities, their action chains, and semantic gaps between language and vision. This…
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
Tuyen Tran, Thao Minh Le, Truyen Tran
Referring-based Video Object Segmentation is a multimodal problem that requires producing fine-grained segmentation results guided by external cues. Traditional approaches to this…
Finding the Trigger: Causal Abductive Reasoning on Video Events
Thao Minh Le, Vuong Le, Kien Do +3
This paper introduces a new problem, Causal Abductive Reasoning on Video Events (CARVE), which involves identifying causal relationships between events in a video and generating hy…
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
Quang-Hung Le, Long Hoang Dang, Ngan Le +2
Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between ent…
Unified Framework with Consistency across Modalities for Human Activity Recognition
Tuyen Tran, Thao Minh Le, Hung Tran +1
Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input m…