7 papers
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
Katsuya Ogata, Zongshang Pang, Mayu Otani +1
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a…
DiverXplorer: Stock Image Exploration via Diversity Adjustment for Graphic Design
Antonio Tejero-de-Pablos, Sichao Song, Naoto Ohsaka +2
Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' abil…
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
Shunki Uebayashi, Kento Masui, Kyohei Atarashi +5
Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their abil…
Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
Zongshang Pang, Mayu Otani, Yuta Nakashima
Temporally localizing user-queried events through natural language is a crucial capability for video models. Recent methods predominantly adapt video LLMs to generate event boundar…
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…
Multimodal Markup Document Models for Graphic Design Completion
Kotaro Kikuchi, Ukyo Honda, Naoto Inoue +3
We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike…