24 citations · 61 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Yi-Fan Zhang, Qingsong Wen, Chaoyou Fu +4
Seeing clearly with high resolution is a foundation of Large Multimodal Models (LMMs), which has been proven to be vital for visual perception and reasoning. Existing works usually…
cs.CV2024
Debiasing Multimodal Large Language Models via Penalization of Language Priors
YiFan Zhang, Yang Shi, Weichen Yu +6
In the realms of computer vision and natural language processing, Multimodal Large Language Models (MLLMs) have become indispensable tools, proficient in generating textual respons…