Action Recognition using Visual Attention
arXiv:1511.04119
Abstract
We propose a soft attention based model for the task of action recognition in videos. We use multi-layered Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units which are deep both spatially and temporally. Our model learns to focus selectively on parts of the video frames and classifies videos after taking a few glimpses. The model essentially learns which parts in the frames are relevant for the task at hand and attaches higher importance to them. We evaluate the model on UCF-11 (YouTube Action), HMDB-51 and Hollywood2 datasets and analyze how the model focuses its attention depending on the scene and the action being performed.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Recurrent Neural Network Regularization
- Theano: new features and speed improvements
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Deep Image: Scaling up Image Recognition
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
Cited by in corpus (108)
- Attention Mechanisms in Computer Vision: A Survey
- Deep neural network models for computational histopathology: A survey
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- Richly Activated Graph Convolutional Network for Robust Skeleton-based Action Recognition
- AUTSL: A Large Scale Multi-modal Turkish Sign Language Dataset and Baseline Methods
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- TSM: Temporal Shift Module for Efficient Video Understanding
- Recurrent Mixture Density Network for Spatiotemporal Visual Attention
- MoVi: A Large Multipurpose Motion and Video Dataset
- Flow-Guided Feature Aggregation for Video Object Detection
- Object Level Visual Reasoning in Videos
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- Infrared and 3D skeleton feature fusion for RGB-D action recognition
- Multi-level Attention Model for Weakly Supervised Audio Classification
- Tensor-Train Recurrent Neural Networks for Video Classification
- DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Impression Network for Video Object Detection
- DeepSignals: Predicting Intent of Drivers Through Visual Signals
- Kronecker CP Decomposition with Fast Multiplication for Compressing RNNs
- Compressing 3DCNNs Based on Tensor Train Decomposition
- Scaling Video Analytics on Constrained Edge Nodes
- Learning to predict where to look in interactive environments using deep recurrent q-learning
- Attention Based Glaucoma Detection: A Large-scale Database and CNN Model
- Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition
- PS-DeVCEM: Pathology-sensitive deep learning model for video capsule endoscopy based on weakly labeled data
- Lattice Long Short-Term Memory for Human Action Recognition
- Call Attention to Rumors: Deep Attention Based Recurrent Neural Networks for Early Rumor Detection
- Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
- Pose-aware Multi-level Feature Network for Human Object Interaction Detection
- Neural Person Search Machines
- Soft + Hardwired Attention: An LSTM Framework for Human Trajectory Prediction and Abnormal Event Detection
- Towards Privacy-Preserving Visual Recognition via Adversarial Training: A Pilot Study
- Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention
- HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval
- Learning Compact Recurrent Neural Networks with Block-Term Tensor Decomposition
- The Benefit of Distraction: Denoising Remote Vitals Measurements using Inverse Attention
- Dynamic Spatial-Temporal Representation Learning for Traffic Flow Prediction
- Cross Domain Knowledge Transfer for Person Re-identification
- Actor-Centric Relation Network
- Understanding More about Human and Machine Attention in Deep Neural Networks
- CSVideoNet: A Real-time End-to-end Learning Framework for High-frame-rate Video Compressive Sensing
- Understanding the computational demands underlying visual reasoning
- Temporal Context Aggregation Network for Temporal Action Proposal Refinement
- Planning on the fast lane: Learning to interact using attention mechanisms in path integral inverse reinforcement learning
- Learning Representative Temporal Features for Action Recognition
- Understanding Visual Ads by Aligning Symbols and Objects using Co-Attention
- Using Spatial Pooler of Hierarchical Temporal Memory to classify noisy videos with predefined complexity
- FoveaNet: Perspective-aware Urban Scene Parsing
- Learning Where to Focus for Efficient Video Object Detection
- Interpretable Spatio-temporal Attention for Video Action Recognition
- Collaborative Attention Mechanism for Multi-View Action Recognition
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Adversarial Cross-Domain Action Recognition with Co-Attention
- Adding Attentiveness to the Neurons in Recurrent Neural Networks
- TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection
- Semantic Object Parsing with Graph LSTM
- Color-wise Attention Network for Low-light Image Enhancement
- CAR-Net: Clairvoyant Attentive Recurrent Network
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Body Joint guided 3D Deep Convolutional Descriptors for Action Recognition
- Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition
- A Unified Method for First and Third Person Action Recognition
- GestARLite: An On-Device Pointing Finger Based Gestural Interface for Smartphones and Video See-Through Head-Mounts
- Person Re-identification Using Visual Attention
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification
- SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
- Action Classification and Highlighting in Videos
- Hierarchical Multi-scale Attention Networks for Action Recognition
- On Attention Modules for Audio-Visual Synchronization
- Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
- A spatiotemporal model with visual attention for video classification
- Spatiotemporal Pyramid Network for Video Action Recognition
- Manipulation-skill Assessment from Videos with Spatial Attention Network
- Spatio-Temporal Instance Learning: Action Tubes from Class Supervision
- Dynamic Filtering with Large Sampling Field for ConvNets
- Reasoning About Human-Object Interactions Through Dual Attention Networks
- Recognizing Video Events with Varying Rhythms
- Two-Stream Video Classification with Cross-Modality Attention
- Fine-grained Video Categorization with Redundancy Reduction Attention
- Temporal Attention-Gated Model for Robust Sequence Classification
- Class Semantics-based Attention for Action Detection
- Action Recognition with Spatio-Temporal Visual Attention on Skeleton Image Sequences
- VideoLightFormer: Lightweight Action Recognition using Transformers
- Temporal-Spatial Mapping for Action Recognition
- Squeeze-and-Excitation on Spatial and Temporal Deep Feature Space for Action Recognition
- Understanding Patch-Based Learning by Explaining Predictions
- FPAN: Fine-grained and Progressive Attention Localization Network for Data Retrieval
- LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
- Memory-Augmented Temporal Dynamic Learning for Action Recognition
- Tri-axial Self-Attention for Concurrent Activity Recognition
- ST-ABN: Visual Explanation Taking into Account Spatio-temporal Information for Video Recognition
- SmallBigNet: Integrating Core and Contextual Views for Video Classification
- Attention-Driven Body Pose Encoding for Human Activity Recognition
- Video-Based Convolutional Attention for Person Re-Identification
- Predictive Coding Networks Meet Action Recognition
- Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks
- Scene Flow to Action Map: A New Representation for RGB-D based Action Recognition with Convolutional Neural Networks
- Higher-order Network for Action Recognition
- Block-term Tensor Neural Networks
- Video-based Person Re-Identification using Gated Convolutional Recurrent Neural Networks
- Object-ABN: Learning to Generate Sharp Attention Maps for Action Recognition
- Learning Object Permanence from Video