activity
20152022
most citedThe Kinetics Human Action Video Dataset

2.9k citations · 3.2k across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL20223 cited

Distribution Aware Metrics for Conditional Natural Language Generation

David M Chan, Yiming Ni, David A Ross +3

Traditional automated metrics for evaluating conditional natural language generation use pairwise comparisons between a single generated text and the best-matching gold-standard gr…

cs.CV2020

Active Learning for Video Description With Cluster-Regularized Ensemble Ranking

David M. Chan, Sudheendra Vijayanarasimhan, David A. Ross +1

Automatic video captioning aims to train models to generate text descriptions for all segments in a video, however, the most effective approaches require large amounts of manual an…

cs.CV2018

Rethinking the Faster R-CNN Architecture for Temporal Action Localization

Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold +3

We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key short…

cs.RO201738 cited

End-to-End Learning of Semantic Grasping

Eric Jang, Sudheendra Vijayanarasimhan, Peter Pastor +2

We consider the task of semantic robotic grasping, in which a robot picks up an object of a user-specified class using only monocular images. Inspired by the two-stream hypothesis…

cs.CV20172.9k cited

The Kinetics Human Action Video Dataset

Will Kay, Joao Carreira, Karen Simonyan +9

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 1…

cs.CV2017

Motion Prediction Under Multimodality with Conditional Stochastic Networks

Katerina Fragkiadaki, Jonathan Huang, Alex Alemi +3

Given a visual history, multiple future outcomes for a video scene are equally probable, in other words, the distribution of future outcomes has multiple modes. Multimodality is no…