The Kinetics Human Action Video Dataset
arXiv:1705.06950
Abstract
We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions are human focussed and cover a broad range of classes including human-object interactions such as playing instruments, as well as human-human interactions such as shaking hands. We describe the statistics of the dataset, how it was collected, and give some baseline performance figures for neural network architectures trained and tested for human action classification on this dataset. We also carry out a preliminary analysis of whether imbalance in the dataset leads to bias in the classifiers.
References in corpus (2)
Cited by in corpus (573)
- Transformers in Vision: A Survey
- GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- DINOv2: Learning Robust Visual Features without Supervision
- Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
- Human Action Recognition from Various Data Modalities: A Review
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- Med3D: Transfer Learning for 3D Medical Image Analysis
- A Short Note about Kinetics-600
- Self-supervised Co-training for Video Representation Learning
- Attention Bottlenecks for Multimodal Fusion
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos
- Learning Video Representations using Contrastive Bidirectional Transformer
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Revisiting ResNets: Improved Training and Scaling Strategies
- Attentional Pooling for Action Recognition
- Patch-VQ: 'Patching Up' the Video Quality Problem
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Benchmarking Micro-action Recognition: Dataset, Methods, and Applications
- TSM: Temporal Shift Module for Efficient Video Understanding
- Anomaly Detection-Inspired Few-Shot Medical Image Segmentation Through Self-Supervision With Supervoxels
- Differentiable Learning-to-Normalize via Switchable Normalization
- Audiovisual SlowFast Networks for Video Recognition
- Non-local Neural Networks
- Self-Supervised Learning for Videos: A Survey
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution
- VideoBERT: A Joint Model for Video and Language Representation Learning
- MDMMT: Multidomain Multimodal Transformer for Video Retrieval
- Adversarial Video Generation on Complex Datasets
- SlowFast Networks for Video Recognition
- Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
- A Comprehensive Study of Deep Video Action Recognition
- Revisiting Temporal Modeling for Video-based Person ReID
- Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
- Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework
- Review: Deep Learning in Electron Microscopy
- ROAD: The ROad event Awareness Dataset for Autonomous Driving
- A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation
- Motion-driven Visual Tempo Learning for Video-based Action Recognition
- TSA-Net: Tube Self-Attention Network for Action Quality Assessment
- RegionViT: Regional-to-Local Attention for Vision Transformers
- Real-time monitoring of driver drowsiness on mobile platforms using 3D neural networks
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
- Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Video Classification with Channel-Separated Convolutional Networks
- The AVA-Kinetics Localized Human Actions Video Dataset
- Spatiotemporal Contrastive Video Representation Learning
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Labelling unlabelled videos from scratch with multi-modal self-supervision
- Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Analyzing Human-Human Interactions: A Survey
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- Using Motion History Images with 3D Convolutional Networks in Isolated Sign Language Recognition
- Deepfake Detection using Spatiotemporal Convolutional Networks
- Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
- Mutual Context Network for Jointly Estimating Egocentric Gaze and Actions
- Space-time Mixing Attention for Video Transformer
- Self-supervised Learning for Video Correspondence Flow
- Continuous Human Action Recognition for Human-Machine Interaction: A Review
- An Image is Worth 16x16 Words, What is a Video Worth?
- Multiscale Vision Transformers
- From CNNs to Transformers in Multimodal Human Action Recognition: A Survey
- Zero-Shot Action Recognition in Videos: A Survey
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Weakly Supervised Action Localization by Sparse Temporal Pooling Network
- The ActivityNet Large-Scale Activity Recognition Challenge 2018 Summary
- Towards Holistic Surgical Scene Understanding
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- Toward Extremely Lightweight Distracted Driver Recognition With Distillation-Based Neural Architecture Search and Knowledge Transfer
- A Multi-viewpoint Outdoor Dataset for Human Action Recognition
- Continual Spatio-Temporal Graph Convolutional Networks
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Video Swin Transformer
- ExtremeWeather: A large-scale climate dataset for semi-supervised detection, localization, and understanding of extreme weather events
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- Video Representation Learning by Dense Predictive Coding
- VideoMix: Rethinking Data Augmentation for Video Classification
- On the effectiveness of task granularity for transfer learning
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- Active Contrastive Learning of Audio-Visual Video Representations
- VARS: Video Assistant Referee System for Automated Soccer Decision Making from Multiple Views
- Appearance-and-Relation Networks for Video Classification
- A Better Baseline for AVA
- ACM-Net: Action Context Modeling Network for Weakly-Supervised Temporal Action Localization
- Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
- Aligning Correlation Information for Domain Adaptation in Action Recognition
- Object Relational Graph with Teacher-Recommended Learning for Video Captioning
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Look, Listen and Learn
- Two-person Graph Convolutional Network for Skeleton-based Human Interaction Recognition
- Shedding Light on Blind Spots: Developing a Reference Architecture to Leverage Video Data for Process Mining
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos
- Watching the World Go By: Representation Learning from Unlabeled Videos
- Anomaly Locality in Video Surveillance
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- Few-shot Action Recognition with Prototype-centered Attentive Learning
- Human Action Performance using Deep Neuro-Fuzzy Recurrent Attention Model
- TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
- Temporal Relational Reasoning in Videos
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- PAN: Towards Fast Action Recognition via Learning Persistence of Appearance
- Cycle-Contrast for Self-Supervised Video Representation Learning
- FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
- TAda! Temporally-Adaptive Convolutions for Video Understanding
- Video Action Understanding
- Natural Environment Benchmarks for Reinforcement Learning
- Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
- OREBA: A Dataset for Objectively Recognizing Eating Behaviour and Associated Intake
- DAVE: A Deep Audio-Visual Embedding for Dynamic Saliency Prediction
- Can Temporal Information Help with Contrastive Self-Supervised Learning?
- STAIR Actions: A Video Dataset of Everyday Home Actions
- TEINet: Towards an Efficient Architecture for Video Recognition
- Slow Motion Matters: A Slow Motion Enhanced Network for Weakly Supervised Temporal Action Localization
- MATIS: Masked-Attention Transformers for Surgical Instrument Segmentation
- G-TAD: Sub-Graph Localization for Temporal Action Detection
- Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art Evaluation
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles
- What Makes Training Multi-Modal Classification Networks Hard?
- Explaining Human Activity Recognition with SHAP: Validating Insights with Perturbation and Quantitative Measures
- Video Understanding as Machine Translation
- Improved Residual Networks for Image and Video Recognition
- Domain Adaptation without Source Data
- Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action Recognition
- Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
- Multimodal Learning for Multi-Omics: A Survey
- TAM: Temporal Adaptive Module for Video Recognition
- Video action detection by learning graph-based spatio-temporal interactions
- Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition
- An Efficient 3D CNN for Action/Object Segmentation in Video
- StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition
- MultiBench: Multiscale Benchmarks for Multimodal Representation Learning
- Learning Representations from Audio-Visual Spatial Alignment
- Sharing Pain: Using Pain Domain Transfer for Video Recognition of Low Grade Orthopedic Pain in Horses
- Action Capsules: Human Skeleton Action Recognition
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog
- Spatio-Temporal FAST 3D Convolutions for Human Action Recognition
- Decompose to manipulate: Manipulable Object Synthesis in 3D Medical Images with Structured Image Decomposition
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection
- Graph Edge Convolutional Neural Networks for Skeleton Based Action Recognition
- A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
- AViD Dataset: Anonymized Videos from Diverse Countries
- Weakly-Supervised Action Localization by Generative Attention Modeling
- All About Knowledge Graphs for Actions
- Hierarchical 3D Feature Learning for Pancreas Segmentation
- Deep Generative Video Compression
- Unsupervised 3D Pose Estimation with Geometric Self-Supervision
- 3D Human Action Representation Learning via Cross-View Consistency Pursuit
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- Study of Subjective and Objective Quality Assessment of Mobile Cloud Gaming Videos
- MIST: Multiple Instance Self-Training Framework for Video Anomaly Detection
- Actor-agnostic Multi-label Action Recognition with Multi-modal Query
- Continual Action Assessment via Task-Consistent Score-Discriminative Feature Distribution Modeling
- Rethinking Motion Representation: Residual Frames with 3D ConvNets for Better Action Recognition
- RGB Stream Is Enough for Temporal Action Detection
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- Fully Transformer-Equipped Architecture for End-to-End Referring Video Object Segmentation
- Temporal Unet: Sample Level Human Action Recognition using WiFi
- TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
- Omni-sourced Webly-supervised Learning for Video Recognition
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- Few-Shot Video Classification via Temporal Alignment
- Memory-Attended Recurrent Network for Video Captioning
- Deep Audio-Visual Learning: A Survey
- Learning Video Representations from Textual Web Supervision
- Graph-Based Global Reasoning Networks
- Temporal Attentive Alignment for Large-Scale Video Domain Adaptation
- GTA: Global Temporal Attention for Video Action Understanding
- Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
- REPAIR: Removing Representation Bias by Dataset Resampling
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- Revisiting 3D ResNets for Video Recognition
- ABN: Agent-Aware Boundary Networks for Temporal Action Proposal Generation
- ViGAT: Bottom-up event recognition and explanation in video using factorized graph attention network
- SimMIM: A Simple Framework for Masked Image Modeling
- Continual 3D Convolutional Neural Networks for Real-time Processing of Videos
- Learning Group Activities from Skeletons without Individual Action Labels
- An Evaluation of Action Recognition Models on EPIC-Kitchens
- Towards Universal Representation for Unseen Action Recognition
- A Closer Look at Few-Shot Video Classification: A New Baseline and Benchmark
- Let's Dance: Learning From Online Dance Videos
- Actor-Centric Relation Network
- -Nets: Double Attention Networks
- Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
- EnsembleNet: End-to-End Optimization of Multi-headed Models
- Action Recognition via Pose-Based Graph Convolutional Networks with Intermediate Dense Supervision
- Multi-Fiber Networks for Video Recognition
- Towards cumulative race time regression in sports: I3D ConvNet transfer learning in ultra-distance running events
- RareAct: A video dataset of unusual interactions
- Surgical Visual Domain Adaptation: Results from the MICCAI 2020 SurgVisDom Challenge
- Deep Analysis of CNN-based Spatio-temporal Representations for Action Recognition
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Variational Cross-Graph Reasoning and Adaptive Structured Semantics Learning for Compositional Temporal Grounding
- Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
- Focusing and Diffusion: Bidirectional Attentive Graph Convolutional Networks for Skeleton-based Action Recognition
- AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- YH Technologies at ActivityNet Challenge 2018
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- ACTION-Net: Multipath Excitation for Action Recognition
- Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue
- Fast Video Shot Transition Localization with Deep Structured Models
- Transformer-based Fusion of 2D-pose and Spatio-temporal Embeddings for Distracted Driver Action Recognition
- Unsupervised Learning from Video with Deep Neural Embeddings
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- The CORSMAL benchmark for the prediction of the properties of containers
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- Spatio-temporal Action Recognition: A Survey
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Listen to Look: Action Recognition by Previewing Audio
- Temporal Sequence Distillation: Towards Few-Frame Action Recognition in Videos
- Model-agnostic Multi-Domain Learning with Domain-Specific Adapters for Action Recognition
- A3D: Adaptive 3D Networks for Video Action Recognition
- Video 3D Sampling for Self-supervised Representation Learning
- Identifying Visible Actions in Lifestyle Vlogs
- Learning to Compress Videos without Computing Motion
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- Segmentations-Leak: Membership Inference Attacks and Defenses in Semantic Image Segmentation
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
- Disentangled Non-Local Neural Networks
- Learning Video Representations from Correspondence Proposals
- PERF-Net: Pose Empowered RGB-Flow Net
- EV-Action: Electromyography-Vision Multi-Modal Action Dataset
- Interactive Fusion of Multi-level Features for Compositional Activity Recognition
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos
- Probing the State of the Art: A Critical Look at Visual Representation Evaluation
- Non-local NetVLAD Encoding for Video Classification
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics
- Real-Time End-to-End Action Detection with Two-Stream Networks
- ElderSim: A Synthetic Data Generation Platform for Human Action Recognition in Eldercare Applications
- Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval
- Black-box Adversarial Attacks on Video Recognition Models
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- ActionFlowNet: Learning Motion Representation for Action Recognition
- Temporal Gaussian Mixture Layer for Videos
- FASTER Recurrent Networks for Efficient Video Classification
- Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
- Adversarial Background-Aware Loss for Weakly-supervised Temporal Activity Localization
- FSD-10: A Dataset for Competitive Sports Content Analysis
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action Recognition
- SSN: Learning Sparse Switchable Normalization via SparsestMax
- Targeted Nonlinear Adversarial Perturbations in Images and Videos
- Trimmed Action Recognition, Dense-Captioning Events in Videos, and Spatio-temporal Action Localization with Focus on ActivityNet Challenge 2019
- Intra- and Inter-Action Understanding via Temporal Action Parsing
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- Understanding the Robustness of Skeleton-based Action Recognition under Adversarial Attack
- 3DTINC: Time-Equivariant Non-Contrastive Learning for Predicting Disease Progression from Longitudinal OCTs
- Bottom-Up Temporal Action Localization with Mutual Regularization
- Learning from Label Relationships in Human Affect
- DeeperForensics Challenge 2020 on Real-World Face Forgery Detection: Methods and Results
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
- Video Self-Stitching Graph Network for Temporal Action Localization
- Self-Attention Network for Skeleton-based Human Action Recognition
- Adversarial Cross-Domain Action Recognition with Co-Attention
- Multi-modal Capsule Routing for Actor and Action Video Segmentation Conditioned on Natural Language Queries
- Towards Training Stronger Video Vision Transformers for EPIC-KITCHENS-100 Action Recognition
- Collaborative Three-Stream Transformers for Video Captioning
- FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
- Reconfigurable Cyber-Physical System for Lifestyle Video-Monitoring via Deep Learning
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- PV-NAS: Practical Neural Architecture Search for Video Recognition
- Relational Action Forecasting
- Group-Skeleton-Based Human Action Recognition in Complex Events
- Uncertainty aware audiovisual activity recognition using deep Bayesian variational inference
- Video-Text Pre-training with Learned Regions
- Malicious or Benign? Towards Effective Content Moderation for Children's Videos
- Density-Guided Label Smoothing for Temporal Localization of Driving Actions
- Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
- Self-supervised Motion Learning from Static Images
- Discriminating Spatial and Temporal Relevance in Deep Taylor Decompositions for Explainable Activity Recognition
- Gated Channel Transformation for Visual Recognition
- Understanding Human Hands in Contact at Internet Scale
- Improving Automated COVID-19 Grading with Convolutional Neural Networks in Computed Tomography Scans: An Ablation Study
- Towards High-Quality Temporal Action Detection with Sparse Proposals
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Content-based Analysis of the Cultural Differences between TikTok and Douyin
- MVFNet: Multi-View Fusion Network for Efficient Video Recognition
- BASAR:Black-box Attack on Skeletal Action Recognition
- Learning to Discretely Compose Reasoning Module Networks for Video Captioning
- Good Practices and A Strong Baseline for Traffic Anomaly Detection
- VideoSSL: Semi-Supervised Learning for Video Classification
- ClusterFit: Improving Generalization of Visual Representations
- Self-supervised Feature Learning by Cross-modality and Cross-view Correspondences
- Volterra Neural Networks (VNNs)
- Temporal Shuffling for Defending Deep Action Recognition Models against Adversarial Attacks
- Three Branches: Detecting Actions With Richer Features
- TinyAction Challenge: Recognizing Real-world Low-resolution Activities in Videos
- Domain Adversarial Reinforcement Learning
- Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
- Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
- Class-Aware Adversarial Lung Nodule Synthesis in CT Images
- Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
- An X3D Neural Network Analysis for Runner's Performance Assessment in a Wild Sporting Environment
- A Large-Scale Re-identification Analysis in Sporting Scenarios: the Betrayal of Reaching a Critical Point
- BAR: Bayesian Activity Recognition using variational inference
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
- Diagnosing Error in Temporal Action Detectors
- Towards Balanced Active Learning for Multimodal Classification
- Federated Action Recognition on Heterogeneous Embedded Devices
- SMART: Skeletal Motion Action Recognition aTtack
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer
- Pose Refinement Graph Convolutional Network for Skeleton-based Action Recognition
- AirObject: A Temporally Evolving Graph Embedding for Object Identification
- Simple Video Generation using Neural ODEs
- Video-based Person Re-identification via 3D Convolutional Networks and Non-local Attention
- RGB-D Based Action Recognition with Light-weight 3D Convolutional Networks
- Fully Automated Hand Hygiene Monitoring\\in Operating Room using 3D Convolutional Neural Network
- Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment
- "Train one, Classify one, Teach one" -- Cross-surgery transfer learning for surgical step recognition
- Pillar Networks++: Distributed non-parametric deep and wide networks
- What can human minimal videos tell us about dynamic recognition models?
- Grounded Objects and Interactions for Video Captioning
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Compositional Few-Shot Recognition with Primitive Discovery and Enhancing
- Zero-Shot Learning with Knowledge Enhanced Visual Semantic Embeddings
- A Multigrid Method for Efficiently Training Video Models
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Low-light Environment Neural Surveillance
- Graph Convolution with Low-rank Learnable Local Filters
- From Detection to Action Recognition: An Edge-Based Pipeline for Robot Human Perception
- 3DPalsyNet: A Facial Palsy Grading and Motion Recognition Framework using Fully 3D Convolutional Neural Networks
- Leveraging Activity Recognition to Enable Protective Behavior Detection in Continuous Data
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Temporal Query Networks for Fine-grained Video Understanding
- BlanketSet -- A clinical real-world in-bed action recognition and qualitative semi-synchronised MoCap dataset
- ImageNet-21K Pretraining for the Masses
- Leveraging Random Label Memorization for Unsupervised Pre-Training
- Learning Efficient Video Representation with Video Shuffle Networks
- Video Action Recognition Via Neural Architecture Searching
- Learning from Videos with Deep Convolutional LSTM Networks
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- STEP: Spatio-Temporal Progressive Learning for Video Action Detection
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- Over-the-Air Adversarial Flickering Attacks against Video Recognition Networks
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- Online Learnable Keyframe Extraction in Videos and its Application with Semantic Word Vector in Action Recognition
- Discriminability Distillation in Group Representation Learning
- Adaptive Focus for Efficient Video Recognition
- Context-Aware RCNN: A Baseline for Action Detection in Videos
- Application of Transfer Learning to Sign Language Recognition using an Inflated 3D Deep Convolutional Neural Network
- Towards Improving Spatiotemporal Action Recognition in Videos
- Spatial-Temporal Alignment Network for Action Recognition and Detection
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
- SAFCAR: Structured Attention Fusion for Compositional Action Recognition
- AVD: Adversarial Video Distillation
- Weakly-supervised Fingerspelling Recognition in British Sign Language Videos
- GPRAR: Graph Convolutional Network based Pose Reconstruction and Action Recognition for Human Trajectory Prediction
- Towards an Unequivocal Representation of Actions
- Learning to Segment Actions from Observation and Narration
- TSI: Temporal Saliency Integration for Video Action Recognition
- Counting Out Time: Class Agnostic Video Repetition Counting in the Wild
- Universal-to-Specific Framework for Complex Action Recognition
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information Bottleneck
- Alleviating Over-segmentation Errors by Detecting Action Boundaries
- Creating a Large-scale Synthetic Dataset for Human Activity Recognition
- Predictively Encoded Graph Convolutional Network for Noise-Robust Skeleton-based Action Recognition
- Temporal Extension Module for Skeleton-Based Action Recognition
- Learning Class Regularized Features for Action Recognition
- Human Action Sequence Classification
- Reasoning About Human-Object Interactions Through Dual Attention Networks
- Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks
- Gradient Weighted Superpixels for Interpretability in CNNs
- Dynamic Inference: A New Approach Toward Efficient Video Action Recognition
- Deep Learning-based Concept Detection in vitrivr at the Video Browser Showdown 2019 - Final Notes
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos
- Iterative Projection and Matching: Finding Structure-preserving Representatives and Its Application to Computer Vision
- Temporal Bilinear Networks for Video Action Recognition
- Pre-training without Natural Images
- Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering
- Skeleton-Based Action Recognition with Synchronous Local and Non-local Spatio-temporal Learning and Frequency Attention
- Random Temporal Skipping for Multirate Video Analysis
- Fine-grained Video Categorization with Redundancy Reduction Attention
- Efficient Action Recognition Using Confidence Distillation
- The Multi-Modal Video Reasoning and Analyzing Competition
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Actional-Structural Graph Convolutional Networks for Skeleton-based Action Recognition
- Multi-Level Temporal Pyramid Network for Action Detection
- Survey: Transformer based Video-Language Pre-training
- Spatio-temporal Video Re-localization by Warp LSTM
- VideoDG: Generalizing Temporal Relations in Videos to Novel Domains
- Unsupervised Multi-label Dataset Generation from Web Data
- Qiniu Submission to ActivityNet Challenge 2018
- Mining for meaning: from vision to language through multiple networks consensus
- Exploiting Inter-Frame Regional Correlation for Efficient Action Recognition
- Massively Parallel Video Networks
- Efficient Video Transformers with Spatial-Temporal Token Selection
- Detecting Kissing Scenes in a Database of Hollywood Films
- UniDual: A Unified Model for Image and Video Understanding
- Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval
- PolyViT: Co-training Vision Transformers on Images, Videos and Audio
- Sequence-to-Sequence Modeling for Action Identification at High Temporal Resolution
- CatNet: Class Incremental 3D ConvNets for Lifelong Egocentric Gesture Recognition
- TimeGate: Conditional Gating of Segments in Long-range Activities
- Few-shot Action Recognition with Implicit Temporal Alignment and Pair Similarity Optimization
- Context-Dependent Models for Predicting and Characterizing Facial Expressiveness
- AIR-Act2Act: Human-human interaction dataset for teaching non-verbal social behaviors to robots
- A Multi-level Alignment Training Scheme for Video-and-Language Grounding
- Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal Data
- Non-local Recurrent Neural Memory for Supervised Sequence Modeling
- Multi-Label Activity Recognition using Activity-specific Features and Activity Correlations
- Adversarial Seeded Sequence Growing for Weakly-Supervised Temporal Action Localization
- Long-Short Temporal Contrastive Learning of Video Transformers
- Two-Stream Video Classification with Cross-Modality Attention
- Effective Action Recognition with Embedded Key Point Shifts
- Visual Concept Reasoning Networks
- Video Modeling with Correlation Networks
- IntegralAction: Pose-driven Feature Integration for Robust Human Action Recognition in Videos
- Temporal Alignment Prediction for Few-Shot Video Classification
- HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
- 2nd Place Scheme on Action Recognition Track of ECCV 2020 VIPriors Challenges: An Efficient Optical Flow Stream Guided Framework
- What's in the Flow? Exploiting Temporal Motion Cues for Unsupervised Generic Event Boundary Detection
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
- An Evaluation of Large Pre-Trained Models for Gesture Recognition using Synthetic Videos
- Efficient Bitrate Ladder Construction using Transfer Learning and Spatio-Temporal Features
- Cricket stroke extraction: Towards creation of a large-scale cricket actions dataset
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
- Multi Scale Temporal Graph Networks For Skeleton-based Action Recognition
- Weakly Supervised Action Selection Learning in Video
- Energy-based Periodicity Mining with Deep Features for Action Repetition Counting in Unconstrained Videos
- Efficient Spatialtemporal Context Modeling for Action Recognition
- Cross-Modal Progressive Comprehension for Referring Segmentation
- WildGait: Learning Gait Representations from Raw Surveillance Streams
- When Video Classification Meets Incremental Classes
- Hierarchical Video Generation for Complex Data
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- Multi-Task Learning of Generalizable Representations for Video Action Recognition
- Global Context Networks
- Inter-intra Variant Dual Representations forSelf-supervised Video Recognition
- A Comprehensive Study on Temporal Modeling for Online Action Detection
- Coronary Artery Segmentation in Angiographic Videos Using A 3D-2D CE-Net
- Out the Window: A Crowd-Sourced Dataset for Activity Classification in Security Video
- SAR-NAS: Skeleton-based Action Recognition via Neural Architecture Searching
- Improving Skeleton-based Action Recognitionwith Robust Spatial and Temporal Features
- Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization
- Localizing the Common Action Among a Few Videos
- Video Contrastive Learning with Global Context
- Unsupervised View-Invariant Human Posture Representation
- Temporal Action Localization with Variance-Aware Networks
- Flow-edge Guided Video Completion
- Learning Temporal Action Proposals With Fewer Labels
- DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue
- Comparison of Spatiotemporal Networks for Learning Video Related Tasks
- AnimGAN: A Spatiotemporally-Conditioned Generative Adversarial Network for Character Animation
- When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs
- ActBERT: Learning Global-Local Video-Text Representations
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Weakly-Supervised Multi-Person Action Recognition in 360 Videos
- Building BROOK: A Multi-modal and Facial Video Database for Human-Vehicle Interaction Research
- Low Pass Filter for Anti-aliasing in Temporal Action Localization
- Encode the Unseen: Predictive Video Hashing for Scalable Mid-Stream Retrieval
- Self-Supervised Video Representation Learning by Video Incoherence Detection
- W-TALC: Weakly-supervised Temporal Activity Localization and Classification
- Spatio-Temporal Video Representation Learning for AI Based Video Playback Style Prediction
- Spatiotemporal Action Recognition in Restaurant Videos
- A Self Validation Network for Object-Level Human Attention Estimation
- Best Vision Technologies Submission to ActivityNet Challenge 2018-Task: Dense-Captioning Events in Videos
- Detecting Attended Visual Targets in Video
- Balanced Representation Learning for Long-tailed Skeleton-based Action Recognition
- Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily Living
- Recent Progress in Appearance-based Action Recognition
- "Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021
- Learning Energy-based Spatial-Temporal Generative ConvNets for Dynamic Patterns
- Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
- A Spectral Nonlocal Block for Neural Networks
- Learning Representations for Pixel-based Control: What Matters and Why?
- PGT: A Progressive Method for Training Models on Long Videos
- PyTorchVideo: A Deep Learning Library for Video Understanding
- Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
- Generalized Few-Shot Video Classification with Video Retrieval and Feature Generation
- LSTA-Net: Long short-term Spatio-Temporal Aggregation Network for Skeleton-based Action Recognition
- Exploring Temporal Information for Improved Video Understanding
- SmallBigNet: Integrating Core and Contextual Views for Video Classification
- SegCodeNet: Color-Coded Segmentation Masks for Activity Detection from Wearable Cameras
- Sequential Feature Filtering Classifier
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- A Real-time Action Representation with Temporal Encoding and Deep Compression
- A Prospective Study on Sequence-Driven Temporal Sampling and Ego-Motion Compensation for Action Recognition in the EPIC-Kitchens Dataset
- Temporal Accumulative Features for Sign Language Recognition
- Knowledge Fusion Transformers for Video Action Recognition
- Learning a Weakly-Supervised Video Actor-Action Segmentation Model with a Wise Selection
- PNL: Efficient Long-Range Dependencies Extraction with Pyramid Non-Local Module for Action Recognition
- Multi-modal Aggregation for Video Classification
- Interaction Graphs for Object Importance Estimation in On-road Driving Videos
- PANDA: A Gigapixel-level Human-centric Video Dataset
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Two-stream Spatiotemporal Feature for Video QA Task
- O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning
- Indoor Scene Recognition in 3D
- Play Fair: Frame Attributions in Video Models
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Predictive Coding Networks Meet Action Recognition
- UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
- Learnable Sampling 3D Convolution for Video Enhancement and Action Recognition
- Discriminative Video Representation Learning Using Support Vector Classifiers
- Pose-based Body Language Recognition for Emotion and Psychiatric Symptom Interpretation
- 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural Networks
- Object-ABN: Learning to Generate Sharp Attention Maps for Action Recognition
- Image to Video Domain Adaptation Using Web Supervision
- ADNet: Temporal Anomaly Detection in Surveillance Videos
- Deep Multi-Kernel Convolutional LSTM Networks and an Attention-Based Mechanism for Videos
- Open Set Domain Adaptation for Image and Action Recognition
- Submission to ActivityNet Challenge 2019: Task B Spatio-temporal Action Localization
- Improving Visual Recognition using Ambient Sound for Supervision
- Few-Shot Transformation of Common Actions into Time and Space
- Zero-Shot Activity Recognition with Videos
- Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
- Are Accelerometers for Activity Recognition a Dead-end?
- CTM: Collaborative Temporal Modeling for Action Recognition
- PLAN-B: Predicting Likely Alternative Next Best Sequences for Action Prediction
- Learning Comprehensive Motion Representation for Action Recognition
- Initialization Using Perlin Noise for Training Networks with a Limited Amount of Data
- Empowering cyberphysical systems of systems with intelligence
- Delta Sampling R-BERT for limited data and low-light action recognition
- RFC-HyPGCN: A Runtime Sparse Feature Compress Accelerator for Skeleton-Based GCNs Action Recognition Model with Hybrid Pruning
- Using phase instead of optical flow for action recognition
- Multi-Object Tracking with Hallucinated and Unlabeled Videos
- Channel-Temporal Attention for First-Person Video Domain Adaptation
- Self Supervision to Distillation for Long-Tailed Visual Recognition
- Interactive Video Retrieval with Dialog
- TraMNet - Transition Matrix Network for Efficient Action Tube Proposals
- Long Short View Feature Decomposition via Contrastive Video Representation Learning
- Few-Shot Adaptation for Multimedia Semantic Indexing
- Disarranged Zone Learning (DZL): An unsupervised and dynamic automatic stenosis recognition methodology based on coronary angiography
- Efficient Modelling Across Time of Human Actions and Interactions
- On Flow Profile Image for Video Representation
- TiVGAN: Text to Image to Video Generation with Step-by-Step Evolutionary Generator
- GTM: Gray Temporal Model for Video Recognition
- The role of ego vision in view-invariant action recognition
- Layout-induced Video Representation for Recognizing Agent-in-Place Actions
- T-RECS: Training for Rate-Invariant Embeddings by Controlling Speed for Action Recognition
- Recurrent Residual Module for Fast Inference in Videos
- Skeleton-Split Framework using Spatial Temporal Graph Convolutional Networks for Action Recogntion
- Multi-object tracking with self-supervised associating network
- Egok360: A 360 Egocentric Kinetic Human Activity Video Dataset
- Hierarchical Action Classification with Network Pruning
- Learning to Sort Image Sequences via Accumulated Temporal Differences
- Approximated Bilinear Modules for Temporal Modeling
- Self-Attentive 3D Human Pose and Shape Estimation from Videos