7 citations · 29 across the 9 of their papers we have counts for
22 papers · 1 filter
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning
AJ Piergiovanni, Weicheng Kuo, Anelia Angelova
We present a simple approach which can turn a ViT encoder into an efficient video model, which can seamlessly work with both image and video inputs. By sparsely sampling the inputs…
Compound Tokens: Channel Fusion for Vision-Language Representation Learning
Maxwell Mbabilla Aladago, AJ Piergiovanni
We present an effective method for fusing visual-and-language representations for several question answering tasks including visual question answering and visual entailment. In con…
Pre-training image-language transformers for open-vocabulary tasks
AJ Piergiovanni, Weicheng Kuo, Anelia Angelova
We present a pre-training approach for vision and language transformer models, which is based on a mixture of diverse tasks. We explore both the use of image-text captioning data i…
4D-Net for Learned Multi-Modal Alignment
AJ Piergiovanni, Vincent Casser, Michael S. Ryoo +1
We present 4D-Net, a 3D object detection approach, which utilizes 3D Point Cloud and RGB sensing information, both in time. We are able to incorporate the 4D information by perform…
Unsupervised Discovery of Actions in Instructional Videos
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities a…
Unsupervised Action Segmentation for Instructional Videos
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. W…