2 papers
cs.CV2025
Generating Accurate and Detailed Captions for High-Resolution Images
Hankyeol Lee, Gawon Seo, Kyounggyu Lee +3
Vision-language models (VLMs) often struggle to generate accurate and detailed captions for high-resolution images since they are typically pre-trained on low-resolution inputs (e.…
cs.CV2024
Enhancing Visual Classification using Comparative Descriptors
Hankyeol Lee, Gawon Seo, Wonseok Choi +3
The performance of vision-language models (VLMs), such as CLIP, in visual classification tasks, has been enhanced by leveraging semantic knowledge from large language models (LLMs)…