Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Visual Instruction Tuning with Chain of Region-of-Interest
Yixin Chen, Shuai Zhang, Boran Han +1
High-resolution (HR) images are pivotal for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs). However, directly increasing image…
cs.CV2024
CaMML: Context-Aware Multimodal Learner for Large Models
Yixin Chen, Shuai Zhang, Boran Han +2
In this work, we introduce Context-Aware MultiModal Learner (CaMML), for tuning large multimodal models (LMMs). CaMML, a lightweight module, is crafted to seamlessly integrate mult…