Non-local Neural Networks
arXiv:1711.07971
Abstract
Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time. In this paper, we present non-local operations as a generic family of building blocks for capturing long-range dependencies. Inspired by the classical non-local means method in computer vision, our non-local operation computes the response at a position as a weighted sum of the features at all positions. This building block can be plugged into many computer vision architectures. On the task of video classification, even without any bells and whistles, our non-local models can compete or outperform current competition winners on both Kinetics and Charades datasets. In static image recognition, our non-local models improve object detection/segmentation and pose estimation on the COCO suite of tasks. Code is available at https://github.com/facebookresearch/video-nonlocal-net .
CVPR 2018, code is available at: https://github.com/facebookresearch/video-nonlocal-net
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improving neural networks by preventing co-adaptation of feature detectors
- Two-Stream Convolutional Networks for Action Recognition in Videos
- WaveNet: A Generative Model for Raw Audio
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- The Kinetics Human Action Video Dataset
- Interaction Networks for Learning about Objects, Relations and Physics
- Spatiotemporal Residual Networks for Video Action Recognition
- Fully Connected Deep Structured Networks
- Visual Interaction Networks
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Image denoising with multi-layer perceptrons, part 2: training trade-offs and analysis of their mechanisms
Cited by in corpus (66)
- Attention U-Net: Learning Where to Look for the Pancreas
- Non-Local Recurrent Network for Image Restoration
- Deformable ConvNets v2: More Deformable, Better Results
- TSM: Temporal Shift Module for Efficient Video Understanding
- Relational Deep Reinforcement Learning
- SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Compact Generalized Non-local Network
- SCAN: Self-and-Collaborative Attention Network for Video Person Re-identification
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Large-Scale Long-Tailed Recognition in an Open World
- Hyperbolic Attention Networks
- Context-self contrastive pretraining for crop type semantic segmentation
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Spatial-Temporal Relation Networks for Multi-Object Tracking
- Relational recurrent neural networks
- Assessing YOLACT++ for real time and robust instance segmentation of medical instruments in endoscopic procedures
- StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
- Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition
- Attacks on State-of-the-Art Face Recognition using Attentional Adversarial Attack Generative Network
- CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
- Knowledge Adaptation for Efficient Semantic Segmentation
- Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition
- Self-Attention Capsule Networks for Object Classification
- PeerNets: Exploiting Peer Wisdom Against Adversarial Attacks
- SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
- Learning to Focus: Cascaded Feature Matching Network for Few-shot Image Recognition
- Graph-Based Global Reasoning Networks
- AdaFrame: Adaptive Frame Selection for Fast Video Recognition
- GTA: Global Temporal Attention for Video Action Understanding
- Triple consistency loss for pairing distributions in GAN-based face synthesis
- Spectral Feature Transformation for Person Re-identification
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- Adaptively Connected Neural Networks
- Deep Exemplar-based Video Colorization
- Hyperspectral Image Classification With Context-Aware Dynamic Graph Convolutional Network
- Learning Visual Question Answering by Bootstrapping Hard Attention
- Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- S2RMs: Spatially Structured Recurrent Modules
- Relational Action Forecasting
- PV-NAS: Practical Neural Architecture Search for Video Recognition
- An Attention-Aided Deep Learning Framework for Massive MIMO Channel Estimation
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Diagnosing Error in Temporal Action Detectors
- Levels of Analysis for Machine Learning
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
- Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
- Local Temporal Bilinear Pooling for Fine-grained Action Parsing
- Video Modeling with Correlation Networks
- Channel Interaction Networks for Fine-Grained Image Categorization
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- GINet: Graph Interaction Network for Scene Parsing
- Fine-grained Image-to-Image Transformation towards Visual Recognition
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Image to Video Domain Adaptation Using Web Supervision
- Cardiac Motion Scoring with Segment- and Subject-level Non-Local Modeling
- Submission to ActivityNet Challenge 2019: Task B Spatio-temporal Action Localization
- Contextualized Spatial-Temporal Network for Taxi Origin-Destination Demand Prediction
- ParNet: Position-aware Aggregated Relation Network for Image-Text matching
- PSGAN++: Robust Detail-Preserving Makeup Transfer and Removal
- Non-local RoIs for Instance Segmentation
- Segmentation of turbulent computational fluid dynamics simulations with unsupervised ensemble learning
- Cyclic orthogonal convolutions for long-range integration of features