activity
20162023
most citedSparse R-CNN: End-to-End Object Detection with Learnable Proposals

103 citations · 311 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

24 papers · 1 filter

cs.CV2023★ 1 cited

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Yan Zeng, Hanbo Zhang, Jiani Zheng +5

Recent advancements in Large Language Models (LLMs) such as GPT4 have displayed exceptional multi-modal capabilities in following open-ended instructions given images. However, the…

cs.CV2023

ClickSeg: 3D Instance Segmentation with Click-Level Weak Annotations

Leyao Liu, Tao Kong, Minzhao Zhu +2

3D instance segmentation methods often require fully-annotated dense labels for training, which are costly to obtain. In this paper, we present ClickSeg, a novel click-level weakly…

cs.CV2022

Towards Unifying Reference Expression Generation and Comprehension

Duo Zheng, Tao Kong, Ya Jing +2

Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a prom…

cs.CV2022★ 10 cited

Generative Category-Level Shape and Pose Estimation with Semantic Primitives

Guanglin Li, Yifeng Li, Zhichao Ye +4

Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for o…

cs.CV2022★ 16 cited

Exploring Target Representations for Masked Autoencoders

Xingbin Liu, Jinghao Zhou, Tao Kong +2

Masked autoencoders have become popular training paradigms for self-supervised visual representation learning. These models randomly mask a portion of the input and reconstruct the…

cs.CV2021

iBOT: Image BERT Pre-Training with Online Tokenizer

Jinghao Zhou, Chen Wei, Huiyu Wang +4

The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pie…