5 papers
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
Hoang-Bao Le, Aiden Durrant, Thai Son Mai +3
Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content and natural language. However,…
Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
Xuekang Zhu, Ji-Zhe Zhou, Kaiwen Feng +5
With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipulation process, comprising a serie…
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
Hoang-Bao Le, Allie Tran, Binh T. Nguyen +2
Image-Guided Retrieval with Optional Text (IGROT) is a general retrieval setting where a query consists of an anchor image, with or without accompanying text, aiming to retrieve se…
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
Hoang-Bao Le, Allie Tran, Binh T. Nguyen +2
Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Imag…
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
Dinh Viet Cuong, Hoang-Bao Le, An Pham Ngoc Nguyen +2
This paper addresses two main objectives. Firstly, we demonstrate the impressive performance of the LLaVA-NeXT-interleave on 22 datasets across three different tasks: Multi-Image R…