5 citations · 7 across the 5 of their papers we have counts for
8 papers · 1 filter
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Minseok Kang, Minhyeok Lee, Jungho Lee +6
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across…
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra +4
Women's safety and security are paramount for a modern society. Often, crimes scenes get recorded through low-resolution CCTV cameras limiting the efficiency of video anomaly detec…
Effective SAM Combination for Open-Vocabulary Semantic Segmentation
Minhyeok Lee, Suhwan Cho, Jungho Lee +4
Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting…
Synchronizing Vision and Language: Bidirectional Token-Masking AutoEncoder for Referring Image Segmentation
Minhyeok Lee, Dogyoon Lee, Jungho Lee +4
Referring Image Segmentation (RIS) aims to segment target objects expressed in natural language within a scene at the pixel level. Various recent RIS models have achieved state-of-…
MAIR: Multi-view Attention Inverse Rendering with 3D Spatially-Varying Lighting Estimation
JunYong Choi, SeokYeong Lee, Haesol Park +3
We propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-vi…
DyAnNet: A Scene Dynamicity Guided Self-Trained Video Anomaly Detection Network
Kamalakar Thakare, Yash Raghuwanshi, Debi Prosad Dogra +2
Unsupervised approaches for video anomaly detection may not perform as good as supervised approaches. However, learning unknown types of anomalies using an unsupervised approach is…