activity
20162024
most citedSocratic Models: Composing Zero-Shot Multimodal Reasoning with Language

173 citations · 234 across the 13 of their papers we have counts for

collaborators
Showing 2018Show all

9 papers · 1 filter

cs.CV2018

Evolving Space-Time Neural Architectures for Videos

AJ Piergiovanni, Anelia Angelova, Alexander Toshev +1

We present a new method for finding video CNN architectures that capture rich spatio-temporal information in videos. Previous work, taking advantage of 3D convolutions, obtained pr…

cs.CV2018

Representation Flow for Action Recognition

AJ Piergiovanni, Michael S. Ryoo

In this paper, we propose a convolutional layer inspired by optical flow algorithms to learn motion representations. Our representation flow layer is a fully-differentiable layer d…

cs.CV2018

Learning Multimodal Representations for Unseen Activities

AJ Piergiovanni, Michael S. Ryoo

We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constra…

cs.RO2018

Learning Real-World Robot Policies by Dreaming

AJ Piergiovanni, Alan Wu, Michael S. Ryoo

Learning to control robots directly based on images is a primary challenge in robotics. However, many existing reinforcement learning approaches require iteratively obtaining milli…

cs.CV2018

Fine-grained Activity Recognition in Baseball Videos

AJ Piergiovanni, Michael S. Ryoo

In this paper, we introduce a challenging new dataset, MLB-YouTube, designed for fine-grained activity detection. The dataset contains two settings: segmented video classification…

cs.CV2018

Musical Chair: Efficient Real-Time Recognition Using Collaborative IoT Devices

Ramyad Hadidi, Jiashen Cao, Matthew Woodward +2

The prevalence of Internet of things (IoT) devices and abundance of sensor data has created an increase in real-time data processing such as recognition of speech, image, and video…