9 papers
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim +2
Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Visi…
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
Sena Korkut, MarÃa Alejandra Bravo Sarmiento, Sanghwan Kim +1
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers shoul…
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Sanghwan Kim, Rui Xiao, Stephan Alaniz +2
Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long v…
Instance-Optimal Estimation with Multiple LLM Judges on a Budget
Junghyun Lee, Sanghwa Kim, Yassir Jedra +2
Evaluating large language models increasingly relies on LLM-as-a-judge protocols, but such evaluations remain costly: different judges have different prices and reliabilities, and…
SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating
Hanbyel Cho, Sang-Hun Kim, Jeonguk Kang +1
Recent advances in real-time interactive text-driven motion generation have enabled humanoids to perform diverse behaviors. However, kinematics-only generators often exhibit physic…
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
Rui Xiao, Sanghwan Kim, Yongqin Xian +2
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coa…