2 papers
cs.CV2026
Visual Reasoning through Tool-supervised Reinforcement Learning
Qihua Dong, Gozde Sahin, Pei Wang +4
In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Models. To achieve that, we pr…
cs.CL2025
Improving Multimodal Large Language Models Using Continual Learning
Shikhar Srivastava, Md Yousuf Harun, Robik Shrestha +1
Generative large language models (LLMs) exhibit impressive capabilities, which can be further augmented by integrating a pre-trained vision model into the original LLM to create a…