9 papers
When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery
Y Huynh, Duc Thanh Nguyen, Thao Minh Le +1
Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to comp…
Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction
Y Huynh, Duc Thanh Nguyen, Mohamed Abdelrazek
The relationship between object perception and reconstruction is well established in human vision, yet remains underexplored in computer vision. In this paper, we demonstrate that…
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
Minh Khoa Le, Kien Do, Duc Thanh Nguyen +1
High-fidelity video generation remains challenging for diffusion models due to the difficulty of modeling complex spatio-temporal dynamics efficiently. Recent video diffusion metho…
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo +4
Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correl…
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
Khanh-Binh Nguyen, Phuoc-Nguyen Bui, Hyunseung Choo +1
Vision-language models (VLMs) exhibit remarkable zero-shot generalization but suffer performance degradation under distribution shifts in downstream tasks, particularly in the abse…
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang +3
Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of und…