1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2025★ 1 cited
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
Zhuoran Yu, Yong Jae Lee
Multimodal Large Language Models (MLLMs) have demonstrated strong performance across a wide range of vision-language tasks, yet their internal processing dynamics remain underexplo…
cs.CV2025
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
Jewon Lee, Ki-Ung Song, Seungmin Yang +6
Visual token reduction lowers inference costs caused by extensive image features in large vision-language models (LVLMs). Unlike relevant studies that prune tokens in self-attentio…