Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment
arXiv:1803.10699
Abstract
In this work, we address the task of weakly-supervised human action segmentation in long, untrimmed videos. Recent methods have relied on expensive learning models, such as Recurrent Neural Networks (RNN) and Hidden Markov Models (HMM). However, these methods suffer from expensive computational cost, thus are unable to be deployed in large scale. To overcome the limitations, the keys to our design are efficiency and scalability. We propose a novel action modeling framework, which consists of a new temporal convolutional network, named Temporal Convolutional Feature Pyramid Network (TCFPN), for predicting frame-wise action labels, and a novel training strategy for weakly-supervised sequence modeling, named Iterative Soft Boundary Assignment (ISBA), to align action sequences and update the network in an iterative fashion. The proposed framework is evaluated on two benchmark datasets, Breakfast and Hollywood Extended, with four different evaluation metrics. Extensive experimental results show that our methods achieve competitive or superior performance to state-of-the-art methods.
CVPR 2018
References in corpus (8)
- Feature Pyramid Networks for Object Detection
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- TricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation
- Weakly Supervised Action Labeling in Videos Under Ordering Constraints
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- Temporal Convolutional Networks for Action Segmentation and Detection
- Weakly supervised learning of actions from transcripts
- An end-to-end generative framework for video segmentation and recognition
Cited by in corpus (18)
- Learning Video Representations using Contrastive Bidirectional Transformer
- A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation
- ASFormer: Transformer for Action Segmentation
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers
- Coarse to Fine Multi-Resolution Temporal Convolutional Network
- Action Segmentation with Joint Self-Supervised Temporal Domain Adaptation
- Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data
- SF-Net: Single-Frame Supervision for Temporal Action Localization
- D3TW: Discriminative Differentiable Dynamic Time Warping for Weakly Supervised Action Alignment and Segmentation
- Weakly Supervised Energy-Based Learning for Action Segmentation
- On Evaluating Weakly Supervised Action Segmentation Methods
- Alleviating Over-segmentation Errors by Detecting Action Boundaries
- Action Shuffle Alternating Learning for Unsupervised Action Segmentation
- Learning to Segment Actions from Observation and Narration
- Learning Discriminative Prototypes with Dynamic Time Warping
- Discovery-and-Selection: Towards Optimal Multiple Instance Learning for Weakly Supervised Object Detection
- Spatio-Temporal Event Segmentation and Localization for Wildlife Extended Videos
- ADNet: Temporal Anomaly Detection in Surveillance Videos