UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
arXiv:1212.0402
Abstract
We introduce UCF101 which is currently the largest dataset of human actions. It consists of 101 action classes, over 13k clips and 27 hours of video data. The database consists of realistic user uploaded videos containing camera motion and cluttered background. Additionally, we provide baseline action recognition results on this new dataset using standard bag of words approach with overall performance of 44.5%. To the best of our knowledge, UCF101 is currently the most challenging dataset of actions due to its large number of classes, large number of clips and also unconstrained nature of such clips.
Cited by in corpus (581)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Learning Transferable Visual Models From Natural Language Supervision
- Transformers in Vision: A Survey
- Learning to Prompt for Vision-Language Models
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Domain Generalization: A Survey
- DINOv2: Learning Robust Visual Features without Supervision
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Human Action Recognition from Various Data Modalities: A Review
- Towards Good Practices for Very Deep Two-Stream ConvNets
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- WebVision Database: Visual Learning and Understanding from Web Data
- A Review on Deep Learning Techniques for Video Prediction
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Self-supervised Co-training for Video Representation Learning
- Beyond Short Snippets: Deep Networks for Video Classification
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Human Activity Recognition Using Tools of Convolutional Neural Networks: A State of the Art Review, Data Sets, Challenges and Future Prospects
- Attentional Pooling for Action Recognition
- Self-Supervised MultiModal Versatile Networks
- ActionCLIP: A New Paradigm for Video Action Recognition
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- TCLR: Temporal Contrastive Learning for Video Representation
- PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition
- A Pursuit of Temporal Accuracy in General Activity Detection
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
- Knowledge Distillation in Deep Learning and its Applications
- A Comprehensive Study of Deep Video Action Recognition
- Evaluating Two-Stream CNN for Video Classification
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
- Rethinking CNN Models for Audio Classification
- Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation
- ROAD: The ROad event Awareness Dataset for Autonomous Driving
- Efficient Two-Stream Network for Violence Detection Using Separable Convolutional LSTM
- Motion-driven Visual Tempo Learning for Video-based Action Recognition
- Self-Supervised Anomaly Detection in Computer Vision and Beyond: A Survey and Outlook
- Real-time monitoring of driver drowsiness on mobile platforms using 3D neural networks
- TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
- Dynamic Sampling Networks for Efficient Action Recognition in Videos
- Spatiotemporal Contrastive Video Representation Learning
- Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
- Labelling unlabelled videos from scratch with multi-modal self-supervision
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
- CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- Hierarchical Attention Network for Action Recognition in Videos
- MFRNet: A New CNN Architecture for Post-Processing and In-loop Filtering
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Context Understanding in Computer Vision: A Survey
- A Good Image Generator Is What You Need for High-Resolution Video Synthesis
- Infrared and 3D skeleton feature fusion for RGB-D action recognition
- VideoLSTM Convolves, Attends and Flows for Action Recognition
- Video Representation Learning with Visual Tempo Consistency
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Multi-level Second-order Few-shot Learning
- An Image is Worth 16x16 Words, What is a Video Worth?
- Label Efficient Learning of Transferable Representations across Domains and Tasks
- MERLOT: Multimodal Neural Script Knowledge Models
- On Learning Sets of Symmetric Elements
- Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
- A Multi-viewpoint Outdoor Dataset for Human Action Recognition
- Temporal Modeling Approaches for Large-scale Youtube-8M Video Understanding
- Video Representation Learning by Dense Predictive Coding
- ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
- VideoMix: Rethinking Data Augmentation for Video Classification
- VARS: Video Assistant Referee System for Automated Soccer Decision Making from Multiple Views
- Active Contrastive Learning of Audio-Visual Video Representations
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation
- Comprehensive Instructional Video Analysis: The COIN Dataset and Performance Evaluation
- CholecTriplet2022: Show me a tool and tell me the triplet -- an endoscopic vision challenge for surgical action triplet detection
- Deep Learning in Physical Layer: Review on Data Driven End-to-End Communication Systems and their Enabling Semantic Applications
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
- Recent Advances in Zero-shot Recognition
- TEA: Temporal Excitation and Aggregation for Action Recognition
- Finding Action Tubes
- Emergence of Exploratory Look-Around Behaviors through Active Observation Completion
- Probabilistic Modeling of Deep Features for Out-of-Distribution and Adversarial Detection
- Few-shot Action Recognition with Prototype-centered Attentive Learning
- Support-set bottlenecks for video-text representation learning
- Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition
- Early Action Prediction with Generative Adversarial Networks
- TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
- VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
- Latent Embedding Feedback and Discriminative Features for Zero-Shot Classification
- Cycle-Contrast for Self-Supervised Video Representation Learning
- PAN: Towards Fast Action Recognition via Learning Persistence of Appearance
- FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- Self-Supervised Visual Learning by Variable Playback Speeds Prediction of a Video
- Hypergraph-based Multi-View Action Recognition using Event Cameras
- Video Action Understanding
- CCVS: Context-aware Controllable Video Synthesis
- Can Temporal Information Help with Contrastive Self-Supervised Learning?
- Learn to cycle: Time-consistent feature discovery for action recognition
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Semi-supervised Body Parsing and Pose Estimation for Enhancing Infant General Movement Assessment
- TEINet: Towards an Efficient Architecture for Video Recognition
- Initialization Strategies of Spatio-Temporal Convolutional Neural Networks
- Unsupervised Representation Learning by Sorting Sequences
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- Complex Sequential Understanding through the Awareness of Spatial and Temporal Concepts
- Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
- Video Understanding as Machine Translation
- Video Imagination from a Single Image with Transformation Generation
- Scaling Laws for Deep Learning
- Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
- Domain Adaptation without Source Data
- Video-based Human Action Recognition using Deep Learning: A Review
- Watch and Learn: Semi-Supervised Learning of Object Detectors from Videos
- DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
- LTC-SUM: Lightweight Client-driven Personalized Video Summarization Framework Using 2D CNN
- StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- Sharing Pain: Using Pain Domain Transfer for Video Recognition of Low Grade Orthopedic Pain in Horses
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- A Review of Emerging Research Directions in Abstract Visual Reasoning
- Learning Representations from Audio-Visual Spatial Alignment
- A Semantic and Motion-Aware Spatiotemporal Transformer Network for Action Detection
- Lattice Long Short-Term Memory for Human Action Recognition
- Grouped Spatial-Temporal Aggregation for Efficient Action Recognition
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
- All About Knowledge Graphs for Actions
- Feature sampling and partitioning for visual vocabulary generation on large action classification datasets
- From Here to There: Video Inbetweening Using Direct 3D Convolutions
- Dual Contrastive Learning for Spatio-temporal Representation
- Muti-view Mouse Social Behaviour Recognition with Deep Graphical Model
- Deep Learning for Vision-based Prediction: A Survey
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
- Rethinking Motion Representation: Residual Frames with 3D ConvNets for Better Action Recognition
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- Learning Discriminative Motion Features Through Detection
- Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution
- Omni-sourced Webly-supervised Learning for Video Recognition
- Music-oriented Dance Video Synthesis with Pose Perceptual Loss
- Recognizing Exercises and Counting Repetitions in Real Time
- InMoDeGAN: Interpretable Motion Decomposition Generative Adversarial Network for Video Generation
- Clean-Label Backdoor Attacks on Video Recognition Models
- Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention
- Local plasticity rules can learn deep representations using self-supervised contrastive predictions
- Copycat CNN: Are Random Non-Labeled Data Enough to Steal Knowledge from Black-box Models?
- VideoPro: A Visual Analytics Approach for Interactive Video Programming
- Unsupervised Learning of Video Representations via Dense Trajectory Clustering
- Few-Shot Video Classification via Temporal Alignment
- ACDnet: An action detection network for real-time edge computing based on flow-guided feature approximation and memory aggregation
- Self-supervised Learning of Audio Representations from Audio-Visual Data using Spatial Alignment
- Learning Video Representations from Textual Web Supervision
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Self-supervised Feature Learning for 3D Medical Images by Playing a Rubik's Cube
- REPAIR: Removing Representation Bias by Dataset Resampling
- Weakly-Supervised Action Localization and Action Recognition using Global-Local Attention of 3D CNN
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph Convolutions
- Biased Mixtures Of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Where and What: Driver Attention-based Object Detection
- Fine-grained Activity Recognition with Holistic and Pose based Features
- Rethinking Full Connectivity in Recurrent Neural Networks
- DNNFusion: Accelerating Deep Neural Networks Execution with Advanced Operator Fusion
- Let's Dance: Learning From Online Dance Videos
- Few-Shot Action Recognition with Compromised Metric via Optimal Transport
- COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
- Action Recognition with Joint Attention on Multi-Level Deep Features
- Neighbor Correspondence Matching for Flow-based Video Frame Synthesis
- BabyNet: A Lightweight Network for Infant Reaching Action Recognition in Unconstrained Environments to Support Future Pediatric Rehabilitation Applications
- Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
- Video Generation from Single Semantic Label Map
- Action Recognition via Pose-Based Graph Convolutional Networks with Intermediate Dense Supervision
- Modeling Multiple Views via Implicitly Preserving Global Consistency and Local Complementarity
- Three-Stream 3D/1D CNN for Fine-Grained Action Classification and Segmentation in Table Tennis
- Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
- Deep Temporal Linear Encoding Networks
- BMBC:Bilateral Motion Estimation with Bilateral Cost Volume for Video Interpolation
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Treatment Learning Causal Transformer for Noisy Image Classification
- Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample
- Spatio-Temporal Fusion Networks for Action Recognition
- Focusing and Diffusion: Bidirectional Attentive Graph Convolutional Networks for Skeleton-based Action Recognition
- Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- Human Action Forecasting by Learning Task Grammars
- Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
- PaStaNet: Toward Human Activity Knowledge Engine
- Temporal Interlacing Network
- Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization
- Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data
- Effect of Architectures and Training Methods on the Performance of Learned Video Frame Prediction
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- ACTION-Net: Multipath Excitation for Action Recognition
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity Recognition
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- Embedded Real-Time Fall Detection Using Deep Learning For Elderly Care
- Learning Person Trajectory Representations for Team Activity Analysis
- Spatio-temporal Action Recognition: A Survey
- Video 3D Sampling for Self-supervised Representation Learning
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
- Model-agnostic Multi-Domain Learning with Domain-Specific Adapters for Action Recognition
- MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned from Image Pairs
- Evidential Deep Learning for Open Set Action Recognition
- The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
- Similarity R-C3D for Few-shot Temporal Activity Detection
- Spatio-Temporal Perturbations for Video Attribution
- Localizing Actions from Video Labels and Pseudo-Annotations
- Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval
- Sparse Adversarial Video Attacks with Spatial Transformations
- Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- Revisiting Few-shot Activity Detection with Class Similarity Control
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild
- Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
- Deep Motion Features for Visual Tracking
- PERF-Net: Pose Empowered RGB-Flow Net
- Black-box Adversarial Attacks on Video Recognition Models
- Pretext-Contrastive Learning: Toward Good Practices in Self-supervised Video Representation Leaning
- Flow-Distilled IP Two-Stream Networks for Compressed Video Action Recognition
- Cross-media Structured Common Space for Multimedia Event Extraction
- Compressed Video Action Recognition with Refined Motion Vector
- Out-of-Distribution Detection for Generalized Zero-Shot Action Recognition
- FSD-10: A Dataset for Competitive Sports Content Analysis
- TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Adversarial Cross-Domain Action Recognition with Co-Attention
- Intra- and Inter-Action Understanding via Temporal Action Parsing
- Two-Stream AMTnet for Action Detection
- Adversarial Background-Aware Loss for Weakly-supervised Temporal Activity Localization
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Conditional Extreme Value Theory for Open Set Video Domain Adaptation
- Progressive Motion Context Refine Network for Efficient Video Frame Interpolation
- Video Is Graph: Structured Graph Module for Video Action Recognition
- Interpretable Deep Feature Propagation for Early Action Recognition
- Relational Action Forecasting
- Self-supervised Motion Learning from Static Images
- FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
- Discriminative convolutional Fisher vector network for action recognition
- GAN for Vision, KG for Relation: a Two-stage Deep Network for Zero-shot Action Recognition
- Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric Transformations
- IF-TTN: Information Fused Temporal Transformation Network for Video Action Recognition
- Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Reconfigurable Cyber-Physical System for Lifestyle Video-Monitoring via Deep Learning
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions
- PV-NAS: Practical Neural Architecture Search for Video Recognition
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Sympathy for the Details: Dense Trajectories and Hybrid Classification Architectures for Action Recognition
- Temporally Coherent Full 3D Mesh Human Pose Recovery from Monocular Video
- Datasets on object manipulation and interaction: a survey
- MAiVAR: Multimodal Audio-Image and Video Action Recognizer
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- A Deep Ranking Model for Spatio-Temporal Highlight Detection from a 360 Video
- Revisiting Rubik's Cube: Self-supervised Learning with Volume-wise Transformation for 3D Medical Image Segmentation
- Learning Dynamic Generator Model by Alternating Back-Propagation Through Time
- Unsupervised Bi-directional Flow-based Video Generation from one Snapshot
- What and How Well You Performed? A Multitask Learning Approach to Action Quality Assessment
- The Best of Both Worlds: Combining Data-independent and Data-driven Approaches for Action Recognition
- Making CNNs for Video Parsing Accessible
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- Encoding Video and Label Priors for Multi-label Video Classification on YouTube-8M dataset
- Eigen Evolution Pooling for Human Action Recognition
- Asynchronous Interaction Aggregation for Action Detection
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Deep Multimodal Feature Encoding for Video Ordering
- Multi-Level Recurrent Residual Networks for Action Recognition
- CDFI: Compression-Driven Network Design for Frame Interpolation
- RoVISQ: Reduction of Video Service Quality via Adversarial Attacks on Deep Learning-based Video Compression
- Towards Structured Analysis of Broadcast Badminton Videos
- Self-Supervised Learning via multi-Transformation Classification for Action Recognition
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- MVFNet: Multi-View Fusion Network for Efficient Video Recognition
- TinyAction Challenge: Recognizing Real-world Low-resolution Activities in Videos
- Audio-visual Representation Learning for Anomaly Events Detection in Crowds
- SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
- Federated Action Recognition on Heterogeneous Embedded Devices
- Searching Action Proposals via Spatial Actionness Estimation and Temporal Path Inference and Tracking
- LP-3DCNN: Unveiling Local Phase in 3D Convolutional Neural Networks
- KIT MOMA: A Mobile Machines Dataset
- fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
- Low-light Environment Neural Surveillance
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
- Busy-Quiet Video Disentangling for Video Classification
- Depthwise Spatio-Temporal STFT Convolutional Neural Networks for Human Action Recognition
- Back to the Future: Cycle Encoding Prediction for Self-supervised Contrastive Video Representation Learning
- Fusing Motion Patterns and Key Visual Information for Semantic Event Recognition in Basketball Videos
- Fully Automated Hand Hygiene Monitoring\\in Operating Room using 3D Convolutional Neural Network
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Human Action Recognition using Local Two-Stream Convolution Neural Network Features and Support Vector Machines
- Enhanced Quadratic Video Interpolation
- Pillar Networks++: Distributed non-parametric deep and wide networks
- Predictive Learning: Using Future Representation Learning Variantial Autoencoder for Human Action Prediction
- Self-supervised Video Representation Learning by Context and Motion Decoupling
- Detection of Object Throwing Behavior in Surveillance Videos
- Multi-shot Temporal Event Localization: a Benchmark
- Action Classification and Highlighting in Videos
- Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment
- Learning Gating ConvNet for Two-Stream based Methods in Action Recognition
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- Detection of Audio-Video Synchronization Errors Via Event Detection
- What can human minimal videos tell us about dynamic recognition models?
- Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Learning from Videos with Deep Convolutional LSTM Networks
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Manipulation-skill Assessment from Videos with Spatial Attention Network
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- From Detection to Action Recognition: An Edge-Based Pipeline for Robot Human Perception
- SPACE: A Simulator for Physical Interactions and Causal Learning in 3D Environments
- Asymmetric Bilateral Motion Estimation for Video Frame Interpolation
- Event and Activity Recognition in Video Surveillance for Cyber-Physical Systems
- Elaborative Rehearsal for Zero-shot Action Recognition
- Spatiotemporal Pyramid Network for Video Action Recognition
- Context-aware and Scale-insensitive Temporal Repetition Counting
- Stacked dense optical flows and dropout layers to predict sperm motility and morphology
- Dense Relational Image Captioning via Multi-task Triple-Stream Networks
- A spatiotemporal model with visual attention for video classification
- Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
- Space-Time-Aware Multi-Resolution Video Enhancement
- Temporal Query Networks for Fine-grained Video Understanding
- Embedding Task Knowledge into 3D Neural Networks via Self-supervised Learning
- Self-Supervised Representation Learning for Visual Anomaly Detection
- Generalized Zero-Shot Learning for Action Recognition with Web-Scale Video Data
- TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
- Alternative Semantic Representations for Zero-Shot Human Action Recognition
- Large-Scale YouTube-8M Video Understanding with Deep Neural Networks
- Cross-Class Relevance Learning for Temporal Concept Localization
- Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
- STH: Spatio-Temporal Hybrid Convolution for Efficient Action Recognition
- Context-Aware RCNN: A Baseline for Action Detection in Videos
- Symmetry and Group in Attribute-Object Compositions
- Single Image Action Recognition by Predicting Space-Time Saliency
- Action Class Relation Detection and Classification Across Multiple Video Datasets
- Global Semantic Descriptors for Zero-Shot Action Recognition
- Residual Frames with Efficient Pseudo-3D CNN for Human Action Recognition
- On the Pitfalls of Learning with Limited Data: A Facial Expression Recognition Case Study
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- Video Action Recognition Via Neural Architecture Searching
- VidTr: Video Transformer Without Convolutions
- Constraint Solving with Deep Learning for Symbolic Execution
- Efficient Video Classification Using Fewer Frames
- ConvGRU in Fine-grained Pitching Action Recognition for Action Outcome Prediction
- Synthetic Defocus and Look-Ahead Autofocus for Casual Videography
- Unified Generator-Classifier for Efficient Zero-Shot Learning
- Mixture of Pre-processing Experts Model for Noise Robust Deep Learning on Resource Constrained Platforms
- Vision-based Behavioral Recognition of Novelty Preference in Pigs
- STEP: Spatio-Temporal Progressive Learning for Video Action Detection
- Compositional Kronecker Context Optimization for Vision-Language Models
- Learning Temporal Embeddings for Complex Video Analysis
- Pooled Motion Features for First-Person Videos
- Video retrieval based on deep convolutional neural network
- Attention Transfer from Web Images for Video Recognition
- Disentangling Motion, Foreground and Background Features in Videos
- AVD: Adversarial Video Distillation
- Temporal Attentive Alignment for Video Domain Adaptation
- Cross-Modal Message Passing for Two-stream Fusion
- Re-Identification Supervised Texture Generation
- Dance with Flow: Two-in-One Stream Action Detection
- Progress Regression RNN for Online Spatial-Temporal Action Localization in Unconstrained Videos
- Better Guider Predicts Future Better: Difference Guided Generative Adversarial Networks
- Technical Report on Visual Quality Assessment for Frame Interpolation
- GradMix: Multi-source Transfer across Domains and Tasks
- Dynamic Inference: A New Approach Toward Efficient Video Action Recognition
- Recognizing Video Events with Varying Rhythms
- Predicting the Future: A Jointly Learnt Model for Action Anticipation
- Scheduled Differentiable Architecture Search for Visual Recognition
- Future Frame Prediction of a Video Sequence
- Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition
- Directional Temporal Modeling for Action Recognition
- Social Adaptive Module for Weakly-supervised Group Activity Recognition
- Universal-to-Specific Framework for Complex Action Recognition
- Robust Visual Object Tracking with Two-Stream Residual Convolutional Networks
- Prediction-assistant Frame Super-Resolution for Video Streaming
- Removing Class Imbalance using Polarity-GAN: An Uncertainty Sampling Approach
- Spatial-Temporal Alignment Network for Action Recognition and Detection
- Depth-Aware Action Recognition: Pose-Motion Encoding through Temporal Heatmaps
- KORSAL: Key-point Detection based Online Real-Time Spatio-Temporal Action Localization
- Learning Cross-modal Contrastive Features for Video Domain Adaptation
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
- Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
- TSI: Temporal Saliency Integration for Video Action Recognition
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile Devices
- Towards Visually Explaining Video Understanding Networks with Perturbation
- TTPP: Temporal Transformer with Progressive Prediction for Efficient Action Anticipation
- Learning Class Regularized Features for Action Recognition
- 2nd Place Scheme on Action Recognition Track of ECCV 2020 VIPriors Challenges: An Efficient Optical Flow Stream Guided Framework
- Learning spatio-temporal representations with temporal squeeze pooling
- Evolution-Preserving Dense Trajectory Descriptors
- Explainable 3D Convolutional Neural Networks by Learning Temporal Transformations
- Detecting Human-to-Human-or-Object (H2O) Interactions with DIABOLO
- Compositional Structure Learning for Action Understanding
- From Actions to Events: A Transfer Learning Approach Using Improved Deep Belief Networks
- Efficient On-the-fly Category Retrieval using ConvNets and GPUs
- Efficient Action Recognition Using Confidence Distillation
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Targeted Attack for Deep Hashing based Retrieval
- Repetitive Activity Counting by Sight and Sound
- TimeGate: Conditional Gating of Segments in Long-range Activities
- CatNet: Class Incremental 3D ConvNets for Lifelong Egocentric Gesture Recognition
- Video Frame Interpolation Transformer
- LIGAR: Lightweight General-purpose Action Recognition
- Exploiting Inter-Frame Regional Correlation for Efficient Action Recognition
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Survey: Transformer based Video-Language Pre-training
- Understanding the Perceived Quality of Video Predictions
- Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- Design Light-weight 3D Convolutional Networks for Video Recognition Temporal Residual, Fully Separable Block, and Fast Algorithm
- Spatio-temporal Video Re-localization by Warp LSTM
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- Human Action Recognition with Deep Temporal Pyramids
- Multi-Label Activity Recognition using Activity-specific Features and Activity Correlations
- Diverse Video Generation using a Gaussian Process Trigger
- Knowledge Integration Networks for Action Recognition
- Identity Preserve Transform: Understand What Activity Classification Models Have Learnt
- Convolutional Neural Network on Three Orthogonal Planes for Dynamic Texture Classification
- Wide and Narrow: Video Prediction from Context and Motion
- CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation
- Crowd Video Captioning
- Few-shot Action Recognition with Implicit Temporal Alignment and Pair Similarity Optimization
- An Experimentation Platform for Explainable Coalition Situational Understanding
- Latent Neural Differential Equations for Video Generation
- Semi-Supervised Few-Shot Atomic Action Recognition
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- Finding a Needle in a Haystack: Tiny Flying Object Detection in 4K Videos using a Joint Detection-and-Tracking Approach
- DFPN: Deformable Frame Prediction Network
- EA-Net: Edge-Aware Network for Flow-based Video Frame Interpolation
- Multi-Field De-interlacing using Deformable Convolution Residual Blocks and Self-Attention
- iMiGUE: An Identity-free Video Dataset for Micro-Gesture Understanding and Emotion Analysis
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- From Traditional to Modern : Domain Adaptation for Action Classification in Short Social Video Clips
- Effective Action Recognition with Embedded Key Point Shifts
- Reinforcement Learning with Latent Flow
- Contrastive Learning of Global-Local Video Representations
- Few Shot Activity Recognition Using Variational Inference
- Personalizing Pre-trained Models
- Token Shift Transformer for Video Classification
- A multimodal deep learning framework for scalable content based visual media retrieval
- Self-Supervised Video Representation Learning by Video Incoherence Detection
- LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
- Recent Progress in Appearance-based Action Recognition
- Batch Normalization with Enhanced Linear Transformation
- Deep set conditioned latent representations for action recognition
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Temporal Bilinear Encoding Network of Audio-Visual Features at Low Sampling Rates
- Machine-Generated Hierarchical Structure of Human Activities to Reveal How Machines Think
- GCF-Net: Gated Clip Fusion Network for Video Action Recognition
- Efficient data-driven encoding of scene motion using Eccentricity
- Time and Frequency Network for Human Action Detection in Videos
- PDWN: Pyramid Deformable Warping Network for Video Interpolation
- One to Transfer All: A Universal Transfer Framework for Vision Foundation Model with Few Data
- Efficient Spatialtemporal Context Modeling for Action Recognition
- AIM 2019 Challenge on Video Temporal Super-Resolution: Methods and Results
- ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
- Video Frame Interpolation via Structure-Motion based Iterative Fusion
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
- Group Activity Prediction with Sequential Relational Anticipation Model
- Recurrent Deconvolutional Generative Adversarial Networks with Application to Text Guided Video Generation
- Three-stream network for enriched Action Recognition
- DeepLandscape: Adversarial Modeling of Landscape Video
- Spatiotemporal Action Recognition in Restaurant Videos
- Temporal Action Localization with Variance-Aware Networks
- Learning Temporal Action Proposals With Fewer Labels
- Towards human-agent knowledge fusion (HAKF) in support of distributed coalition teams
- CS-MCNet:A Video Compressive Sensing Reconstruction Network with Interpretable Motion Compensation
- A Spectral Nonlocal Block for Neural Networks
- Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition
- Improvements of Motion Estimation and Coding using Neural Networks
- UAV-GESTURE: A Dataset for UAV Control and Gesture Recognition
- ODN: Opening the Deep Network for Open-set Action Recognition
- HRVGAN: High Resolution Video Generation using Spatio-Temporal GAN
- Video Summarization via Actionness Ranking
- Semantic Adversarial Network with Multi-scale Pyramid Attention for Video Classification
- Human Activity Recognition for Edge Devices
- Ontology Based Global and Collective Motion Patterns for Event Classification in Basketball Videos
- Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-supervised Regression
- Generalized Few-Shot Video Classification with Video Retrieval and Feature Generation
- Articulated motion discovery using pairs of trajectories
- Memory-Augmented Temporal Dynamic Learning for Action Recognition
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Deep Spatio-temporal Manifold Network for Action Recognition
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- Joint Max Margin and Semantic Features for Continuous Event Detection in Complex Scenes
- Multi-kernel learning of deep convolutional features for action recognition
- Dimensionality Reduction on Grassmannian via Riemannian Optimization: A Generalized Perspective
- Tubelets: Unsupervised action proposals from spatiotemporal super-voxels
- Channel Pruning Guided by Classification Loss and Feature Importance
- Inter-intra Variant Dual Representations forSelf-supervised Video Recognition
- Video Contrastive Learning with Global Context
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition
- Egok360: A 360 Egocentric Kinetic Human Activity Video Dataset
- Pose-based Body Language Recognition for Emotion and Psychiatric Symptom Interpretation
- Deep Sequence Learning for Video Anticipation: From Discrete and Deterministic to Continuous and Stochastic
- Large-Scale Mapping of Human Activity using Geo-Tagged Videos
- A stepped sampling method for video detection using LSTM
- A Proposed Artificial intelligence Model for Real-Time Human Action Localization and Tracking
- Skeleton-Split Framework using Spatial Temporal Graph Convolutional Networks for Action Recogntion
- FREGAN : an application of generative adversarial networks in enhancing the frame rate of videos
- Channel-Temporal Attention for First-Person Video Domain Adaptation
- Advancing Video Self-Supervised Learning via Image Foundation Models
- Modality Compensation Network: Cross-Modal Adaptation for Action Recognition
- CTM: Collaborative Temporal Modeling for Action Recognition
- GTM: Gray Temporal Model for Video Recognition
- Long-Short Temporal Modeling for Efficient Action Recognition
- Face-Focused Cross-Stream Network for Deception Detection in Videos
- Nuisance-Label Supervision: Robustness Improvement by Free Labels
- Long Short-Term Relation Networks for Video Action Detection
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- PersonRank: Detecting Important People in Images
- Exploring Frame Segmentation Networks for Temporal Action Localization
- Empowering cyberphysical systems of systems with intelligence
- Revisiting hand-crafted feature for action recognition: a set of improved dense trajectories
- Feature-Supervised Action Modality Transfer
- Fusing Deep Convolutional Networks for Large Scale Visual Concept Classification
- Perceptron Synthesis Network: Rethinking the Action Scale Variances in Videos
- Detecting Biological Locomotion in Video: A Computational Approach
- Motion Representation with Acceleration Images
- Multi-View Intact Space Learning
- DDLSTM: Dual-Domain LSTM for Cross-Dataset Action Recognition
- Boosting Video Representation Learning with Multi-Faceted Integration
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-Adaptation
- Learning Single/Multi-Attribute of Object with Symmetry and Group
- On Flow Profile Image for Video Representation
- Deep hierarchical pooling design for cross-granularity action recognition
- PreViTS: Contrastive Pretraining with Video Tracking Supervision
- Few-Shot Transformation of Common Actions into Time and Space
- Exploring Temporal Information for Improved Video Understanding
- PNL: Efficient Long-Range Dependencies Extraction with Pyramid Non-Local Module for Action Recognition
- A Real-time Action Representation with Temporal Encoding and Deep Compression
- CLTA: Contents and Length-based Temporal Attention for Few-shot Action Recognition
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- Exploring Feature Representation and Training strategies in Temporal Action Localization
- Efficient Modelling Across Time of Human Actions and Interactions
- Physical Context and Timing Aware Sequence Generating GANs
- Unsupervised Motion Representation Enhanced Network for Action Recognition
- Approximated Bilinear Modules for Temporal Modeling
- Self-supervised learning using consistency regularization of spatio-temporal data augmentation for action recognition
- MV-C3D: A Spatial Correlated Multi-View 3D Convolutional Neural Networks
- Discovering Multi-Label Actor-Action Association in a Weakly Supervised Setting
- ArrowGAN : Learning to Generate Videos by Learning Arrow of Time
- Action parsing using context features
- SegCodeNet: Color-Coded Segmentation Masks for Activity Detection from Wearable Cameras
- Long Short View Feature Decomposition via Contrastive Video Representation Learning
- CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Gradient Boundary Histograms for Action Recognition
- Delta Sampling R-BERT for limited data and low-light action recognition
- Object-ABN: Learning to Generate Sharp Attention Maps for Action Recognition
- Finding Action Tubes with a Sparse-to-Dense Framework
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- Knowledge Fusion Transformers for Video Action Recognition
- Smoothed Gaussian Mixture Models for Video Classification and Recommendation
- Temporal Accumulative Features for Sign Language Recognition
- ViP: Video Platform for PyTorch
- 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural Networks
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Predictive Coding Networks Meet Action Recognition
- Adaptive Future Frame Prediction with Ensemble Network
- Deep Learning for Fitness
- Learning to Sort Image Sequences via Accumulated Temporal Differences