5 papers
Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models
Meng Zheng, Samhita Marri, Anwesa Choudhuri +6
Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacki…
EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
Van-Loc Nguyen, AprilPyone MaungMaung, Minh-Triet Tran +1
Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its semantic type, and ground its de…
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
Dinh-Khoi Vo, Van-Loc Nguyen, Tam V. Nguyen +2
Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often re…
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran +1
Event-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and…
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Trong-Thuan Nguyen, Viet-Tham Huynh, Quang-Thuc Nguyen +30
Recent 3D retrieval systems are typically designed for simple, controlled scenarios, such as identifying an object from a cropped image or a brief description. However, real-world…