activity
20222024
most citedInternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

17 citations · 25 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20242 cited

Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Chuofan Ma, Yi Jiang, Jiannan Wu +2

We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region…

cs.CV202417 cited

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Zhe Chen, Jiannan Wu, Wenhai Wang +12

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundati…

cs.CV2023

Exploring Transformers for Open-world Instance Segmentation

Jiannan Wu, Yi Jiang, Bin Yan +3

Open-world instance segmentation is a rising task, which aims to segment all objects in the image by learning from a limited number of base-category objects. This task is challengi…

cs.CV20231 cited

Multi-Level Contrastive Learning for Dense Prediction Task

Qiushan Guo, Yizhou Yu, Yi Jiang +3

In this work, we present Multi-Level Contrastive Learning for Dense Prediction Task (MCL), an efficient self-supervised method for learning region-level feature representation for…

cs.CV20225 cited

Language as Queries for Referring Video Object Segmentation

Jiannan Wu, Yi Jiang, Peize Sun +2

Referring video object segmentation (R-VOS) is an emerging cross-modal task that aims to segment the target object referred by a language expression in all video frames. In this wo…