Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
Allen He, Qi Liu, Kun Liu +2
Temporal sentence grounding in videos (TSGV) aims to localize a temporal segment that semantically corresponds to a sentence query from an untrimmed video. Most current methods ado…
cs.CV2024
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
Qi Liu, Xinchen Liu, Kun Liu +2
Multi-modal human action segmentation is a critical and challenging task with a wide range of applications. Nowadays, the majority of approaches concentrate on the fusion of dense…