4 papers
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
Hyungjin Chung, Hyelin Nam, Jiyeon Kim +6
Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail leads to prohibitive computational…
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
Seongsu Ha, Chaeyun Kim, Donghwa Kim +3
Referring Image Segmentation is a comprehensive task to segment an object referred by a textual query from an image. In nature, the level of difficulty in this task is affected by…
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
Junho Lee, Jeongwoo Shin, Seung Woo Ko +2
Given a video with frames, frame sampling is a task to select frames, so as to maximize the performance of a fixed video classifier. Not just brute-force search, but…
TWLV-I: Analysis and Insights from Holistic Evaluation on Video Foundation Models
Hyeongmin Lee, Jin-Young Kim, Kyungjune Baek +18
In this work, we discuss evaluating video foundation models in a fair and robust manner. Unlike language or image foundation models, many video foundation models are evaluated with…