51 citations · 100 across the 6 of their papers we have counts for
8 papers
3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation
Zihao Xiao, Longlong Jing, Shangxuan Wu +9
3D panoptic segmentation is a challenging perception task, especially in autonomous driving. It aims to predict both semantic and instance annotations for 3D points in a scene. Alt…
Contrastive Feature Masking Open-Vocabulary Vision Transformer
Dahun Kim, Anelia Angelova, Weicheng Kuo
We present Contrastive Feature Masking Vision Transformer (CFM-ViT) - an image-text pretraining methodology that achieves simultaneous learning of image- and region-level represent…
Joint Adaptive Representations for Image-Language Learning
AJ Piergiovanni, Anelia Angelova
Image-language learning has made unprecedented progress in visual understanding. These developments have come at high costs, as contemporary vision-language models require large mo…
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Xi Chen, Josip Djolonga, Piotr Padlewski +40
We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…
Video Question Answering with Iterative Video-Text Co-Tokenization
AJ Piergiovanni, Kairo Morton, Weicheng Kuo +2
Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal in…
Mechanical Search on Shelves with Efficient Stacking and Destacking of Objects
Huang Huang, Letian Fu, Michael Danielczuk +6
Stacking increases storage efficiency in shelves, but the lack of visibility and accessibility makes the mechanical search problem of revealing and extracting target objects diffic…