2 papers
cs.CL2026
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
Shunki Uebayashi, Kento Masui, Kyohei Atarashi +5
Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their abil…
cs.CV2025
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
Qianru Qiu, Jiafeng Mao, Kento Masui +1
Recent advances in diffusion models have significantly improved the performance of reference-guided line art colorization. However, existing methods still struggle with region-leve…