3 papers
cs.CV2026
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra +4
Women's safety and security are paramount for a modern society. Often, crimes scenes get recorded through low-resolution CCTV cameras limiting the efficiency of video anomaly detec…
cs.CV2026
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Minseok Kang, Minhyeok Lee, Jungho Lee +6
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across…
cs.CV2025
Effective SAM Combination for Open-Vocabulary Semantic Segmentation
Minhyeok Lee, Suhwan Cho, Jungho Lee +4
Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting…