2 papers
cs.CV2023
Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding
Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed +3
While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length…
cs.CV2022
Unifying Tracking and Image-Video Object Detection
Peirong Liu, Rui Wang, Pengchuan Zhang +6
Objection detection (OD) has been one of the most fundamental tasks in computer vision. Recent developments in deep learning have pushed the performance of image OD to new heights…