5 papers
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Beomsik Cho, Jinhyeong Kim, Dongseok Lee +1
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do…
SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding
Seoyeon Kim, Minjae Kang, Jaehyung Kim
Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts o…
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
Beomsik Cho, Jaehyung Kim
Large Vision Language Models (LVLMs) achieve strong performance across multimodal tasks by integrating visual perception with language understanding. However, how vision informatio…
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
Dongyoung Kim, Sumin Park, Woomin Song +6
Improving embodied reasoning in multimodal-large-language models (MLLMs) is essential for building vision-language-action models (VLAs) on top of them to readily translate multimod…
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
Minjae Kang, Jaehyung Kim
Large Language Models (LLMs), despite advances in instruction tuning, often fail to follow complex user instructions. Activation steering techniques aim to mitigate this by manipul…