Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Delin Mao, Chenghao Sun, Jingwei Song +2
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different…
cs.CV2026
Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
Chengsheng Zhang, Chenghao Sun, Zhining Xie +1
Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understanding. However, the internal m…