most citedRethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2024

Pluralistic Salient Object Detection

Xuelu Feng, Yunsheng Li, Dongdong Chen +4

We introduce pluralistic salient object detection (PSOD), a novel task aimed at generating multiple plausible salient segmentation results for a given input image. Unlike conventio…

cs.CV20241 cited

Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

Yuanze Lin, Yunsheng Li, Dongdong Chen +4

In recent years, multimodal large language models (MLLMs) have made significant strides by training on vast high-quality image-text datasets, enabling them to generally understand…

cs.CV2024

OmniVid: A Generative Framework for Universal Video Understanding

Junke Wang, Dongdong Chen, Chong Luo +4

The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution.…

eess.IV2024

Generative Enhancement for 3D Medical Images

Lingting Zhu, Noel Codella, Dongdong Chen +3

The limited availability of 3D medical image datasets, due to privacy concerns and high collection or annotation costs, poses significant challenges in the field of medical imaging…

cs.CV20231 cited

PersonMAE: Person Re-Identification Pre-Training with Masked AutoEncoders

Hezhen Hu, Xiaoyi Dong, Jianmin Bao +4

Pre-training is playing an increasingly important role in learning generic feature representation for Person Re-identification (ReID). We argue that a high-quality ReID representat…