2 papers
cs.CV2026
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
Katsuya Ogata, Zongshang Pang, Mayu Otani +1
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a…
cs.CV2026
Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
Zongshang Pang, Mayu Otani, Yuta Nakashima
Temporally localizing user-queried events through natural language is a crucial capability for video models. Recent methods predominantly adapt video LLMs to generate event boundar…