7 citations · 24 across the 7 of their papers we have counts for
13 papers · 1 filter
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang +9
Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existi…
Intriguing Properties of Large Language and Vision Models
Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim +2
Recently, large language and vision models (LLVMs) have received significant attention and development efforts due to their remarkable generalization performance across a wide rang…
Group Generalized Mean Pooling for Vision Transformer
Byungsoo Ko, Han-Gyu Kim, Byeongho Heo +4
Vision Transformer (ViT) extracts the final representation from either class token or an average of all patch tokens, following the architecture of Transformer in Natural Language…
Granularity-aware Adaptation for Image Retrieval over Multiple Tasks
Jon Almazán, Byungsoo Ko, Geonmo Gu +2
Strong image search models can be learned for a specific domain, ie. set of labels, provided that some labeled images of that domain are available. A practical visual search model,…
Large-scale Bilingual Language-Image Contrastive Learning
Byungsoo Ko, Geonmo Gu
This paper is a technical report to share our experience and findings building a Korean and English bilingual multimodal model. While many of the multimodal datasets focus on Engli…
RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network
Minchul Shin, Yoonjae Cho, Byungsoo Ko +1
In this paper, we study the compositional learning of images and texts for image retrieval. The query is given in the form of an image and text that describes the desired modificat…