2 papers
cs.CV2026
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Minseok Kang, Minhyeok Lee, Jungho Lee +6
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across…
cs.CV2026
Cross Pseudo Labeling For Weakly Supervised Video Anomaly Detection
Dayeon Lee, Donghyeong Kim, Chaewon Park +2
Weakly supervised video anomaly detection aims to detect anomalies and identify abnormal categories with only video-level labels. We propose CPL-VAD, a dual-branch framework with c…