58 citations · 184 across the 88 of their papers we have counts for
Showing 2024 · cs.CVShow all
3 papers · 2 filters
cs.CV2024
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Qing Jiang, Gen Luo, Yuqin Yang +5
Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they…
cs.CV2024★ 2 cited
LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
Weiquan Huang, Aoqi Wu, Yifan Yang +10
CLIP is a seminal multimodal model that maps images and text into a shared representation space through contrastive learning on billions of image-caption pairs. Inspired by the rap…
cs.CV2024★ 6 cited
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
Xingxuan Zhang, Jiansheng Li, Wenjing Chu +6
We investigate the generalization boundaries of current Multimodal Large Language Models (MLLMs) via comprehensive evaluation under out-of-distribution scenarios and domain-specifi…