most citedMultiModal-GPT: A Vision and Language Model for Dialogue with Humans

65 citations · 99 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV202320 cited

LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Peng Xu, Wenqi Shao, Kaipeng Zhang +7

Large Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their…

cs.CV202365 cited

MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Tao Gong, Chengqi Lyu, Shilong Zhang +7

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generat…

cs.CV20222 cited

Large-batch Optimization for Dense Visual Predictions

Zeyue Xue, Jianming Liang, Guanglu Song +4

Training a large-scale deep neural network in a large-scale dataset is challenging and time-consuming. The recent breakthrough of large-batch optimization is a promising way to tac…

cs.CV202210 cited

Rethinking Resolution in the Context of Efficient Video Recognition

Chuofan Ma, Qiushan Guo, Yi Jiang +3

In this paper, we empirically study how to make the most of low-resolution frames for efficient video recognition. Existing methods mainly focus on developing compact networks or a…

cs.CV2022

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

Yuying Ge, Yixiao Ge, Xihui Liu +5

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast gl…

cs.CV20222 cited

Semantic-Aware Pretraining for Dense Video Captioning

Teng Wang, Zhu Liu, Feng Zheng +3

This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video…