activity
20232025
most citedInternVideo2: Scaling Foundation Models for Multimodal Video Understanding

16 citations · 17 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2025

Seg-VAR: Image Segmentation with Visual Autoregressive Modeling

Rongkun Zheng, Lu Qi, Xi Chen +3

While visual autoregressive modeling (VAR) strategies have shed light on image generation with the autoregressive models, their potential for segmentation, a task that requires pre…

cs.CV2025

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

Jiahe Zhao, Rongkun Zheng, Yi Wang +2

In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input…

cs.CV2024

SyncVIS: Synchronized Video Instance Segmentation

Rongkun Zheng, Lu Qi, Xi Chen +4

Recent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal inform…

cs.CV2024

ViLLa: Video Reasoning Segmentation with Large Language Model

Rongkun Zheng, Lu Qi, Xi Chen +4

Recent efforts in video reasoning segmentation (VRS) integrate large language models (LLMs) with perception models to localize and track objects via textual instructions, achieving…

cs.CV2024★ 16 cited

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Yi Wang, Kunchang Li, Xinhao Li +17

We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialog…

cs.CV2023★ 1 cited

TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

Rongkun Zheng, Lu Qi, Xi Chen +4

Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What…