activity
20182020
most citedAttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

7 citations · 18 across the 4 of their papers we have counts for

collaborators

10 papers

cs.CV20201 cited

AssembleNet++: Assembling Modality Representations via Attention Connections

Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta +1

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy at…

cs.CV2020

Adversarial Generative Grammars for Human Activity Prediction

AJ Piergiovanni, Anelia Angelova, Alexander Toshev +1

In this paper we propose an adversarial generative grammar model for future prediction. The objective is to learn a model that explicitly captures temporal dependencies, providing…

cs.CV20207 cited

AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

Xiaofang Wang, Xuehan Xiong, Maxim Neumann +5

Convolutional operations have two limitations: (1) do not explicitly model where to focus as the same filter is applied to all the positions, and (2) are unsuitable for modeling lo…

cs.CV2020

Evolving Losses for Unsupervised Video Representation Learning

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

We present a new method to learn video representations from large-scale unlabeled video data. Ideally, this representation will be generic and transferable, directly usable for new…

cs.RO20196 cited

Model-based Behavioral Cloning with Future Image Similarity Learning

Alan Wu, AJ Piergiovanni, Michael S. Ryoo

We present a visual imitation learning framework that enables learning of robot action policies solely based on expert samples without any robot trials. Robot exploration and on-po…

cs.CV20194 cited

Evolving Losses for Unlabeled Video Representation Learning

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

We present a new method to learn video representations from unlabeled data. Given large-scale unlabeled video data, the objective is to benefit from such data by learning a generic…