4 papers · 1 filter
ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets
Hoyoung Kim, Minwoo Jang, Jabin Koo +2
Beyond general recognition tasks, specialized domains and fine-grained settings often encounter data scarcity, especially for tail classes. To obtain less biased and more reliable…
Emergence of Text Readability in Vision Language Models
Jaeyoo Park, Sanghyuk Chun, Wonjae Kim +2
We investigate how the ability to recognize textual content within images emerges during the training of Vision-Language Models (VLMs). Our analysis reveals a critical phenomenon:…
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
Sanghyuk Chun, Sangdoo Yun
Recently, Probabilistic Language-Image Pre-Training (ProLIP) has been proposed to tackle the multiplicity issue of vision-language (VL) tasks. Despite their success in probabilisti…
Probabilistic Language-Image Pre-Training
Sanghyuk Chun, Wonjae Kim, Song Park +1
Vision-language models (VLMs) embed aligned image-text pairs into a joint space but often rely on deterministic embeddings, assuming a one-to-one correspondence between images and…