activity
20242026
most citedLLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs

5 citations · 5 across the 10 of their papers we have counts for

collaborators

11 papers

cs.CV2026

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara

The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. T…

cs.CL2026

Neuron Level Analysis of Large Language Model in Legal Domain Reasoning

Eri Onami, Youmi Ma, Shuhei Kurita +1

We presented a neuron-level analysis of legal-domain reasoning in LLMs, comparing it with other applied domain tasks across seven open-weight models. Using neuron attribution score…

cs.CV2026

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

Issa Sugiura, Shuhei Kurita, Yusuke Oda +1

Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchmarks have advanced rapidly, n…

cs.CV2026

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

Daichi Yashima, Shuhei Kurita, Yusuke Oda +3

In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging because of their intricate…

cs.CV2026

JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation

Issa Sugiura, Koki Maeda, Shuhei Kurita +3

Reliable evaluation is essential for the development of vision-language models (VLMs). However, Japanese VQA benchmarks have undergone far less iterative refinement than their Engl…

cs.CV2026

Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

Issa Sugiura, Keito Sasagawa, Keisuke Nakao +8

Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically c…