Large-Scale Mapping of Human Activity using Geo-Tagged Videos
arXiv:1706.07911
Abstract
This paper is the first work to perform spatio-temporal mapping of human activity using the visual content of geo-tagged videos. We utilize a recent deep-learning based video analysis framework, termed hidden two-stream networks, to recognize a range of activities in YouTube videos. This framework is efficient and can run in real time or faster which is important for recognizing events as they occur in streaming video or for reducing latency in analyzing already captured video. This is, in turn, important for using video in smart-city applications. We perform a series of experiments to show our approach is able to accurately map activities both spatially and temporally. We also demonstrate the advantages of using the visual content over the tags/titles.
Accepted at ACM SIGSPATIAL 2017
References in corpus (11)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- YouTube-8M: A Large-Scale Video Classification Benchmark
- FlowNet: Learning Optical Flow with Convolutional Networks
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Urban Magnetism Through The Lens of Geo-tagged Photography
- Back to Basics: Unsupervised Learning of Optical Flow via Brightness Constancy and Motion Smoothness
- Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
- Deep Temporal Linear Encoding Networks