Online Spatiotemporal Action Detection and Prediction via Causal Representations
arXiv:2008.13759
Abstract
In this thesis, we focus on video action understanding problems from an online and real-time processing point of view. We start with the conversion of the traditional offline spatiotemporal action detection pipeline into an online spatiotemporal action tube detection system. An action tube is a set of bounding connected over time, which bounds an action instance in space and time. Next, we explore the future prediction capabilities of such detection methods by extending an existing action tube into the future by regression. Later, we seek to establish that online/causal representations can achieve similar performance to that of offline three dimensional (3D) convolutional neural networks (CNNs) on various tasks, including action recognition, temporal action segmentation and early prediction.
PhD thesis, Oxford Brookes University, Examiners: Dr. Andrea Vedaldi and Dr. Fridolin Wild, 172 pages
References in corpus (9)
- Deep Learning in Neural Networks: An Overview
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge
- AntisymmetricRNN: A Dynamical System View on Recurrent Neural Networks
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Generic Tubelet Proposals for Action Localization