3 papers
cs.CL2026
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability
Kyosuke Takami, Yuka Tateisi, Satoshi Sekine +1
Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japanese K-12 assessments remain…
cs.CL2025
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
Hisami Suzuki, Satoru Katsumata, Takashi Kodama +3
In this paper we present AnswerCarefully, a dataset for promoting the safety and appropriateness of Japanese LLM outputs. The dataset consists of 1,800 pairs of questions and refer…
cs.CL2024
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
LLM-jp, :, Akiko Aizawa +80
This paper introduces LLM-jp, a cross-organizational project for the research and development of Japanese large language models (LLMs). LLM-jp aims to develop open-source and stron…