activity
20222024
most citedGrounded SAM: Assembling Open-World Models for Diverse Visual Tasks

93 citations · 125 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Causal Diffusion Transformers for Generative Modeling

Chaorui Deng, Deyao Zhu, Kunchang Li +2

We introduce Causal Diffusion as the autoregressive (AR) counterpart of Diffusion models. It is a next-token(s) forecasting framework that is friendly to both discrete and continuo…

cs.CV2024

VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model

Xinhao Li, Zhenpeng Huang, Jing Wang +2

With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their r…

cs.CV202493 cited

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Tianhe Ren, Shilong Liu, Ailing Zeng +14

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and seg…

cs.CV2024

Vlogger: Make Your Dream A Vlog

Shaobin Zhuang, Kunchang Li, Xinyuan Chen +4

In this work, we present Vlogger, a generic AI system for generating a minute-level video blog (i.e., vlog) of user descriptions. Different from short videos with a few seconds, vl…

cs.CV202314 cited

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Zhaoyang Liu, Yinan He, Wenhai Wang +17

We present an interactive visual framework named InternGPT, or iGPT for short. The framework integrates chatbots that have planning and reasoning capabilities, such as ChatGPT, wit…

cs.CV202218 cited

Tip-Adapter: Training-free Adaption of CLIP for Few-shot Classification

Renrui Zhang, Zhang Wei, Rongyao Fang +5

Contrastive Vision-Language Pre-training, known as CLIP, has provided a new paradigm for learning visual representations using large-scale image-text pairs. It shows impressive per…