collaborators

10 papers

cs.CV2025

Probabilistic Language-Image Pre-Training

Sanghyuk Chun, Wonjae Kim, Song Park +1

Vision-language models (VLMs) embed aligned image-text pairs into a joint space but often rely on deterministic embeddings, assuming a one-to-one correspondence between images and…

cs.CV2025

Emergence of Text Readability in Vision Language Models

Jaeyoo Park, Sanghyuk Chun, Wonjae Kim +2

We investigate how the ability to recognize textual content within images emerges during the training of Vision-Language Models (VLMs). Our analysis reveals a critical phenomenon:…

cs.CV2025

An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval

Jaeseok Byun, Seokhyeon Jeong, Wonjae Kim +2

Composed Image Retrieval (CIR) aims to retrieve a target image based on a reference image and conditioning text, enabling controllable image searches. The mainstream Zero-Shot (ZS)…

cs.CV2025

LongProLIP: A Probabilistic Vision-Language Model with Long Context Text

Sanghyuk Chun, Sangdoo Yun

Recently, Probabilistic Language-Image Pre-Training (ProLIP) has been proposed to tackle the multiplicity issue of vision-language (VL) tasks. Despite their success in probabilisti…

cs.LG2025

DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias

Song Park, Sanghyuk Chun, Byeongho Heo +1

This paper argues that deep neural networks (DNNs) mostly determine their outputs during the early stages of inference, where biases inherent in the model play a crucial role in sh…

cs.CV2024

RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models

Seulki Park, Daeho Um, Hajung Yoon +3

With the extensive use of vision-language models in various downstream tasks, evaluating their robustness is crucial. In this paper, we propose a benchmark for assessing the robust…