Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Self-Evolving Code-with-Image Reasoning
Tianze Yang, Liang Wu, Ruitong Sun +6
Multimodal models increasingly reach for tools when solving visual tasks (crop, zoom, rotate, brighten), a paradigm known as thinking-with-images. The central challenge is one of p…
cs.CV2026
Self-Improving Small Object Grounding in LVLMs
Tianze Yang, Yucheng Shi, Ruitong Sun +2
Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provide an affirmative answer. At…