8 papers
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao +3
This paper introduces the first systematic evaluation framework for quantifying the quality and risks of papers written by modern coding agents. While AI-driven paper writing has b…
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
Atsuyuki Miyai, Shota Onohara, Jeonghun Baek +1
This paper introduces JMMMU-Pro, an image-based Japanese Multi-discipline Multimodal Understanding Benchmark, and Vibe Benchmark Construction, a scalable construction method. Follo…
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
Atsuyuki Miyai, Mashiro Toyooka, Takashi Otonari +2
Understanding the current capabilities and risks of AI Scientist systems (autoresearch) is essential for ensuring trustworthy and sustainable AI-driven scientific progress while pr…
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
Tatsuki Kawakami, Kazuki Egashira, Atsuyuki Miyai +2
In recent years, unlearning techniques, which are methods for inducing a model to "forget" previously learned information, have attracted attention as a way to address privacy and…
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
Jeonghun Baek, Kazuki Egashira, Shota Onohara +4
Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives…
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
Shiho Noda, Atsuyuki Miyai, Qing Yu +2
Out-of-distribution (OOD) detection is a task that detects OOD samples during inference to ensure the safety of deployed models. However, conventional benchmarks have reached perfo…