2 papers
cs.CV2026
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
Xin Dong, Wenjia Geng, Wenfeng Deng +1
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given…
cs.CV2025
Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
Sujia Wang, Xiangwei Shen, Yansong Tang +3
Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-fram…