103 citations · 311 across the 22 of their papers we have counts for
24 papers · 1 filter
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Yan Zeng, Hanbo Zhang, Jiani Zheng +5
Recent advancements in Large Language Models (LLMs) such as GPT4 have displayed exceptional multi-modal capabilities in following open-ended instructions given images. However, the…
ClickSeg: 3D Instance Segmentation with Click-Level Weak Annotations
Leyao Liu, Tao Kong, Minzhao Zhu +2
3D instance segmentation methods often require fully-annotated dense labels for training, which are costly to obtain. In this paper, we present ClickSeg, a novel click-level weakly…
Towards Unifying Reference Expression Generation and Comprehension
Duo Zheng, Tao Kong, Ya Jing +2
Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a prom…
Generative Category-Level Shape and Pose Estimation with Semantic Primitives
Guanglin Li, Yifeng Li, Zhichao Ye +4
Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for o…
Exploring Target Representations for Masked Autoencoders
Xingbin Liu, Jinghao Zhou, Tao Kong +2
Masked autoencoders have become popular training paradigms for self-supervised visual representation learning. These models randomly mask a portion of the input and reconstruct the…
iBOT: Image BERT Pre-Training with Online Tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang +4
The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pie…