most citedHow Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

16 citations · 24 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV202416 cited

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Zhe Chen, Weiyun Wang, Hao Tian +32

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models…

cs.CV20247 cited

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Xiaoyi Dong, Pan Zhang, Yuhang Zang +21

The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due…

cs.CV20231 cited

PersonMAE: Person Re-Identification Pre-Training with Masked AutoEncoders

Hezhen Hu, Xiaoyi Dong, Jianmin Bao +4

Pre-training is playing an increasingly important role in learning generic feature representation for Person Re-identification (ReID). We argue that a high-quality ReID representat…

cs.GR2023

Emotional Listener Portrait: Neural Listener Head Generation with Emotion

Luchuan Song, Guojun Yin, Zhenchao Jin +2

Listener head generation centers on generating non-verbal behaviors (e.g., smile) of a listener in reference to the information delivered by a speaker. A significant challenge when…

cs.CV2023

Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain Prompting

Qidong Huang, Xiaoyi Dong, Dongdong Chen +5

In this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has…