4 papers
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
Woohyeon Park, Jaeik Kim, Sunghwan Steve Cho +7
Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechani…
A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model
Jihun Park, Jongmin Gim, Kyoungmin Lee +5
We present a training-free style-aligned image generation method that leverages a scale-wise autoregressive model. While large-scale text-to-image (T2I) models, particularly diffus…
MMPB: It's Time for Multi-Modal Personalization
Jaeik Kim, Woojin Kim, Woohyeon Park +1
Visual personalization is essential in user-facing AI systems such as smart homes and healthcare, where aligning model behavior with user-centric concepts is critical. However, rec…
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
Woohyeon Park, Woojin Kim, Jaeik Kim +1
Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accu…