2 papers
cs.AI2026
Reason Through the Latent! Making Latent Visual Reasoning Necessary
Suhyeong Park, Junha Jung, Jaewoo Kang
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being pres…
cs.IR2026
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Suhyeong Park, Junha Jung, Jungwoo Park +1
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expe…