Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Video Token Merging for Long-form Video Understanding
Seon-Ho Lee, Jue Wang, Zhikang Zhang +2
As the scale of data and models for video understanding rapidly expand, handling long-form video input in transformer-based models presents a practical challenge. Rather than resor…
cs.CV2023
MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation
Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee +5
Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal a…