activity
20172025
most citedPaLI: A Jointly-Scaled Multilingual Language-Image Model

196 citations · 378 across the 23 of their papers we have counts for

collaborators
Showing 2018Show all

6 papers · 1 filter

cs.CV2018

Evolving Space-Time Neural Architectures for Videos

AJ Piergiovanni, Anelia Angelova, Alexander Toshev +1

We present a new method for finding video CNN architectures that capture rich spatio-temporal information in videos. Previous work, taking advantage of 3D convolutions, obtained pr…

cs.CV2018

Representation Flow for Action Recognition

AJ Piergiovanni, Michael S. Ryoo

In this paper, we propose a convolutional layer inspired by optical flow algorithms to learn motion representations. Our representation flow layer is a fully-differentiable layer d…

cs.CV2018

Learning Multimodal Representations for Unseen Activities

AJ Piergiovanni, Michael S. Ryoo

We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constra…

cs.RO2018

Learning Real-World Robot Policies by Dreaming

AJ Piergiovanni, Alan Wu, Michael S. Ryoo

Learning to control robots directly based on images is a primary challenge in robotics. However, many existing reinforcement learning approaches require iteratively obtaining milli…

cs.CV2018

Fine-grained Activity Recognition in Baseball Videos

AJ Piergiovanni, Michael S. Ryoo

In this paper, we introduce a challenging new dataset, MLB-YouTube, designed for fine-grained activity detection. The dataset contains two settings: segmented video classification…

cs.CV2018

Temporal Gaussian Mixture Layer for Videos

AJ Piergiovanni, Michael S. Ryoo

We introduce a new convolutional layer named the Temporal Gaussian Mixture (TGM) layer and present how it can be used to efficiently capture longer-term temporal information in con…