most citedLanguage as Queries for Referring Video Object Segmentation

5 citations · 11 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2024

IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model

Yatai Ji, Shilong Zhang, Jie Wu +6

The rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of…

cs.CV20245 cited

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Peize Sun, Yi Jiang, Shoufa Chen +4

We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an…

cs.CV2023

Enhancing Your Trained DETRs with Box Refinement

Yiqun Chen, Qiang Chen, Peize Sun +3

We present a conceptually simple, efficient, and general framework for localization problems in DETR-like models. We add plugins to well-trained models instead of inefficiently des…

cs.CV20231 cited

Going Denser with Open-Vocabulary Part Segmentation

Peize Sun, Shoufa Chen, Chenchen Zhu +4

Object detection has been expanded from a limited number of categories to open vocabulary. Moving forward, a complete intelligent vision system requires understanding more fine-gra…

cs.CV20235 cited

ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box

Yifu Zhang, Xinggang Wang, Xiaoqing Ye +6

Multi-object tracking (MOT) aims at estimating bounding boxes and identities of objects across video frames. Detection boxes serve as the basis of both 2D and 3D MOT. The inevitabl…

cs.CV2022

Towards Grand Unification of Object Tracking

Bin Yan, Yi Jiang, Peize Sun +4

We present a unified method, termed Unicorn, that can simultaneously solve four tracking problems (SOT, MOT, VOS, MOTS) with a single network using the same model parameters. Due t…