Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
arXiv:1502.07209 · doi:10.1109/TPAMI.2017.2670560
Abstract
In this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Although extensive efforts have been devoted in recent years, most existing works combined multiple video features using simple fusion strategies and neglected the utilization of inter-class semantic relationships. This paper proposes a novel unified framework that jointly exploits the feature relationships and the class relationships for improved categorization performance. Specifically, these two types of relationships are estimated and utilized by rigorously imposing regularizations in the learning process of a deep neural network (DNN). Such a regularized DNN (rDNN) can be efficiently realized using a GPU-based implementation with an affordable training cost. Through arming the DNN with better capability of harnessing both the feature and the class relationships, the proposed rDNN is more suitable for modeling video semantics. With extensive experimental evaluations, we show that rDNN produces superior performance over several state-of-the-art approaches. On the well-known Hollywood2 and Columbia Consumer Video benchmarks, we obtain very competitive results: 66.9\% and 73.5\% respectively in terms of mean average precision. In addition, to substantially evaluate our rDNN and stimulate future research on large scale video categorization, we collect and release a new benchmark dataset, called FCVID, which contains 91,223 Internet videos and 239 manually annotated categories.
Please cite the officially published IEEE TPAMI version if you find this work helpful
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Unsupervised Learning of Video Representations using LSTMs
- Multi-Task Feature Learning Via Efficient l2,1-Norm Minimization
- Clustered Multi-Task Learning: A Convex Formulation
Cited by in corpus (56)
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Human Action Recognition from Various Data Modalities: A Review
- Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Creating A Multi-track Classical Musical Performance Dataset for Multimodal Music Analysis: Challenges, Insights, and Applications
- Learning to Recommend with Multiple Cascading Behaviors
- A survey on trajectory clustering analysis
- What Objective Does Self-paced Learning Indeed Optimize?
- Recent Advances in Zero-shot Recognition
- Deep Multimodality Learning for UAV Video Aesthetic Quality Assessment
- Transfer Learning for Video Recognition with Scarce Training Data for Deep Convolutional Neural Network
- A Large-scale Varying-view RGB-D Action Dataset for Arbitrary-view Human Action Recognition
- Clean-Label Backdoor Attacks on Video Recognition Models
- AdaFrame: Adaptive Frame Selection for Fast Video Recognition
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video Hashing
- A Novel Quantum N-Queens Solver Algorithm and its Simulation and Application to Satellite Communication Using IBM Quantum Experience
- ViGAT: Bottom-up event recognition and explanation in video using factorized graph attention network
- EventNet: A Large Scale Structured Concept Library for Complex Event Detection in Video
- Live Laparoscopic Video Retrieval with Compressed Uncertainty
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Query Attack via Opposite-Direction Feature:Towards Robust Image Retrieval
- Black-box Adversarial Attacks on Video Recognition Models
- HCMS: Hierarchical and Conditional Modality Selection for Efficient Video Recognition
- AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and Localization
- Learning to score the figure skating sports videos
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition
- VRFP: On-the-fly Video Retrieval using Web Images and Fast Fisher Vector Products
- Adaptive Focus for Efficient Video Recognition
- FastVA: Deep Learning Video Analytics Through Edge Processing and NPU in Mobile
- Video Imprint
- Dynamic Network Quantization for Efficient Video Inference
- Video Representation Learning and Latent Concept Mining for Large-scale Multi-label Video Classification
- A Proposal-based Approach for Activity Image-to-Video Retrieval
- Towards ontology driven learning of visual concept detectors
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- FlowChroma -- A Deep Recurrent Neural Network for Video Colorization
- Video Stream Retrieval of Unseen Queries using Semantic Memory
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Machine-Generated Hierarchical Structure of Human Activities to Reveal How Machines Think
- Online Action Detection in Streaming Videos with Time Buffers
- Encode the Unseen: Predictive Video Hashing for Scalable Mid-Stream Retrieval
- Multi-Label Zero-Shot Human Action Recognition via Joint Latent Ranking Embedding
- Counting Grid Aggregation for Event Retrieval and Recognition
- Cascaded Coarse-to-Fine Deep Kernel Networks for Efficient Satellite Image Change Detection
- A Case Study of Deep-Learned Activations via Hand-Crafted Audio Features
- Large-Scale Mapping of Human Activity using Geo-Tagged Videos
- Multi-modal Aggregation for Video Classification
- Skimming and Scanning for Untrimmed Video Action Recognition
- Nonlinear Tensor Ring Network
- Anomaly Detection in Video Sequences: A Benchmark and Computational Model
- Large age-gap face verification by feature injection in deep networks
- Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
- Large-Scale Video Classification with Feature Space Augmentation coupled with Learned Label Relations and Ensembling
- Training compact deep learning models for video classification using circulant matrices