5 papers · 1 filter
Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese
Naho Orita, Hayato Ogawa, Daisuke Kawahara
Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Draw…
Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models
Rei Minamoto, Yusuke Oda, Daisuke Kawahara
Sensitive personal information can appear in large-scale pre-training corpora for large language models (LLMs). Detecting and filtering such information is therefore essential to e…
ShapleyLaw: A Game-Theoretic Approach to Multilingual Scaling Laws
Xuyang Cao, Qianying Liu, Chuan Xiao +7
In multilingual pretraining, the test loss of a pretrained model is heavily influenced by the proportion of each language in the pretraining data, namely the \textit{language mixtu…
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
LLM-jp, :, Akiko Aizawa +80
This paper introduces LLM-jp, a cross-organizational project for the research and development of Japanese large language models (LLMs). LLM-jp aims to develop open-source and stron…
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
Keito Sasagawa, Koki Maeda, Issa Sugiura +3
To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While m…