activity
20162026
most citedPermutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2023

Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

Syed Waleed Hyder, Muhammad Usama, Anas Zafar +5

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which dir…

cs.CV2023

Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion

Quoc-Huy Tran, Muhammad Ahmed, Murad Popattia +3

This paper presents a self-supervised temporal video alignment framework which is useful for several fine-grained human activity understanding applications. In contrast with the st…

cs.CV2023★ 1 cited

Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment

Quoc-Huy Tran, Ahmed Mehmood, Muhammad Ahmed +4

This paper presents an unsupervised transformer-based framework for temporal activity segmentation which leverages not only frame-level cues but also segment-level cues. This is in…

cs.CV2022

Timestamp-Supervised Action Segmentation with Graph Convolutional Networks

Hamza Khan, Sanjay Haresh, Awais Ahmed +4

We introduce a novel approach for temporal activity segmentation with timestamp supervision. Our main contribution is a graph convolutional network, which is learned in an end-to-e…

cs.CV2021

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

Sateesh Kumar, Sanjay Haresh, Awais Ahmed +3

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and onlin…

cs.CV2021

Learning by Aligning Videos in Time

Sanjay Haresh, Sateesh Kumar, Huseyin Coskun +4

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level informa…