4 papers
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
Dinh-Khoi Vo, Van-Loc Nguyen, Tam V. Nguyen +2
Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often re…
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran +1
Event-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and…
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Trong-Thuan Nguyen, Viet-Tham Huynh, Quang-Thuc Nguyen +30
Recent 3D retrieval systems are typically designed for simple, controlled scenarios, such as identifying an object from a cropped image or a brief description. However, real-world…
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran +1
Retrieving 3D objects in complex indoor environments using only a masked 2D image and a natural language description presents significant challenges. The ROOMELSA challenge limits…