13 citations · 24 across the 4 of their papers we have counts for
14 papers
UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen +2
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts o…
Exploring Long-Sequence Masked Autoencoders
Ronghang Hu, Shoubhik Debnath, Saining Xie +1
Masked Autoencoding (MAE) has emerged as an effective approach for pre-training representations across multiple domains. In contrast to discrete tokens in natural languages, the in…
UniT: Multimodal Multitask Learning with a Unified Transformer
Ronghang Hu, Amanpreet Singh
We propose UniT, a Unified Transformer model to simultaneously learn the most prominent tasks across different domains, ranging from object detection to natural language understand…
Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single Image
Ronghang Hu, Nikhila Ravi, Alexander C. Berg +1
We present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input…
TextCaps: a Dataset for Image Captioning with Reading Comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1
Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical char…
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
Ronghang Hu, Amanpreet Singh, Trevor Darrell +1
Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water la…