5 papers
Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval
Boseung Jeong, Taegyu Park, Donghyeon Kwon +2
At the heart of composed visual data retrieval is the fusion of a reference visual input and a textual modification into a single query. While current state-of-the-art methods util…
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
Boseung Jeong, Jicheol Park, Sungyeon Kim +1
Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval.…
Improving Text-based Person Search via Part-level Cross-modal Correspondence
Jicheol Park, Boseung Jeong, Dongwon Kim +1
Text-based person search is the task of finding person images that are the most relevant to the natural language text description given as query. The main challenge of this task is…
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
Jicheol Park, Dongwon Kim, Boseung Jeong +1
Text-based person search, employing free-form text queries to identify individuals within a vast image collection, presents a unique challenge in aligning visual and textual repres…
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
Sungyeon Kim, Boseung Jeong, Donghyun Kim +1
Large-scale image-text pre-trained models enable zero-shot classification and provide consistent accuracy across various data distributions. Nonetheless, optimizing these models in…