Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models
Chenyu Huang, Peng Ye, Xudong Tan +4
Efficiently enhancing the reasoning capabilities of Vision-Language Models (VLMs) by merging them with Large Reasoning Models (LRMs) has emerged as a promising direction. However,…
cs.CV2026
Can Multimodal Large Language Models Truly Understand Small Objects?
Fujun Han, Junan Chen, Xintong Zhu +4
Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they re…