6 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CV2026
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception
Chengzhi Mao, Xudong Lin, Wen-Sheng Chu
Vision foundation models are typically trained as static feature extractors, placing the burden of task adaptation onto large downstream models. We propose an alternative paradigm:…
cs.CV2024★ 6 cited
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
Xingxuan Zhang, Jiansheng Li, Wenjing Chu +6
We investigate the generalization boundaries of current Multimodal Large Language Models (MLLMs) via comprehensive evaluation under out-of-distribution scenarios and domain-specifi…