Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Test-Time Training for Visual Foresight Vision-Language-Action Models
Sangwu Park, Wonjoong Kim, Yeonjun In +3
Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inherent design of VF-VLA makes it…
cs.CV2026
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
Jiwan Kim, Kibum Kim, Wonjoong Kim +2
Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that while existing pruning meth…
cs.CV2025
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
Wonjoong Kim, Sangwu Park, Yeonjun In +2
Recently, interpreting complex charts with logical reasoning has emerged as challenges due to the development of vision-language models. A prior state-of-the-art (SOTA) model has p…