activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Early Failure Detection and Intervention in Video Diffusion Models

Kwon Byung-Ki, Sohwi Lim, Nam Hyeon-Woo +2

Text-to-video (T2V) diffusion models have rapidly advanced, yet generations still occasionally fail in practice, such as low text-video alignment or low perceptual quality. Since d…

cs.CV2025

Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching

Wonseok Choi, Sohwi Lim, Nam Hyeon-Woo +4

Instance-level image retrieval aims to find images containing the same object as a given query, despite variations in size, position, or appearance. To address this challenging tas…

cs.CV2025

RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models

Moon Ye-Bin, Roy Miles, Tae-Hyun Oh +2

Image retouching not only enhances visual quality but also serves as a means of expressing personal preferences and emotions. However, existing learning-based approaches require la…

cs.CV2024

VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models

Nam Hyeon-Woo, Moon Ye-Bin, Wonseok Choi +2

Vision language models (VLMs) have shown promising reasoning capabilities across various benchmarks; however, our understanding of their visual perception remains limited. In this…

cs.CV2024

BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models

Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi +1

Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-t…

cs.CV2024

SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems

Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi +3

Data imbalance in training data often leads to biased predictions from trained models, which in turn causes ethical and social issues. A straightforward solution is to carefully cu…