1 citations · 1 across the 3 of their papers we have counts for
4 papers
Selective Test-Time Debiasing for CLIP via Reward Gating
Jaeho Han, Jisoo Yang, Hyeondong Woo +3
Vision language models (VLMs) demonstrate strong zero-shot performance, but often perpetuate social stereotypes in person-centric queries, yielding skewed demographic distributions…
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
Woojun Jung, Junyeong Kim
Video-to-text summarization remains underexplored in terms of comprehensive evaluation methods. Traditional n-gram overlap-based metrics and recent large language model (LLM)-based…
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
Woojun Jung, Jaehoon Go, Mingyu Jeon +2
Multimodal Large Language Models (MLLMs) demonstrate impressive reasoning capabilities, but often fail to perceive fine-grained visual details, limiting their applicability in prec…
See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval
Mingyu Jeon, Sungjin Han, Jinkwon Hwang +3
Recent advances in Multimodal Large Language Models (MLLMs) have improved image recognition and reasoning, but video-related tasks remain challenging due to memory constraints from…