7 papers
DiCE-CIR: Direct Composition Learning for Efficient Zero-Shot Composed Image Retrieval
Gwang-Ho Na, Ho-Joong Kim, Seong-Whan Lee
Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image from a multimodal query consisting of a reference image and an edit text describing the desired modifica…
ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection
Jae-Ryung Hong, Ho-Joong Kim, Seong-Whan Lee
Open-vocabulary object detection aims to localize and classify objects beyond the fixed set of categories seen dur ing training. Recent open-vocabulary object detection methods imp…
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
Ji-Hyeon Kim, Ho-Joong Kim, Seong-Whan Lee
Video moment retrieval is the task of retrieving specific segments of a video corresponding to a given text query. Recent studies have been conducted to improve multimodal alignmen…
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
Ju-Young Oh, Ho-Joong Kim, Seong-Whan Lee
Video question answering (VQA) is a multimodal task that requires the interpretation of a video to answer a given question. Existing VQA methods primarily utilize question and answ…
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
Jung-Ho Hong, Ho-Joong Kim, Kyu-Sung Jeon +1
The feature attribution method reveals the contribution of input variables to the decision-making process to provide an attribution map for explanation. Existing methods grounded o…
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
Yeong-Joon Ju, Ho-Joong Kim, Seong-Whan Lee
Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often direc…