activity
20242026
collaborators

7 papers

cs.CV2026

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

Katsuya Ogata, Zongshang Pang, Mayu Otani +1

Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a…

cs.HC2026

DiverXplorer: Stock Image Exploration via Diversity Adjustment for Graphic Design

Antonio Tejero-de-Pablos, Sichao Song, Naoto Ohsaka +2

Graphic designers explore large stock image collections during open-ended or early-stage design tasks, yet common tools emphasize relevance and similarity, limiting designers' abil…

cs.CL2026

Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory

Shunki Uebayashi, Kento Masui, Kyohei Atarashi +5

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their abil…

cs.CV2026

Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs

Zongshang Pang, Mayu Otani, Yuta Nakashima

Temporally localizing user-queried events through natural language is a crucial capability for video models. Recent methods predominantly adapt video LLMs to generate event boundar…

cs.CV2026

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2025

Multimodal Markup Document Models for Graphic Design Completion

Kotaro Kikuchi, Ukyo Honda, Naoto Inoue +3

We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike…