68 citations · 171 across the 9 of their papers we have counts for
3 papers · 2 filters
UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen +2
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts o…
Scaling Language-Image Pre-training via Masking
Yanghao Li, Haoqi Fan, Ronghang Hu +2
We present Fast Language-Image Pre-training (FLIP), a simple and more efficient method for training CLIP. Our method randomly masks out and removes a large portion of image patches…
Exploring Long-Sequence Masked Autoencoders
Ronghang Hu, Shoubhik Debnath, Saining Xie +1
Masked Autoencoding (MAE) has emerged as an effective approach for pre-training representations across multiple domains. In contrast to discrete tokens in natural languages, the in…