2 papers
cs.CV2026
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar +4
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reas…
cs.LG2026
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Rosie Zhao, Anshul Shah, Xiaoyu Zhu +5
Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-langua…