3 papers
cs.CL2025
PLaMo 2 Technical Report
Preferred Networks, :, Kaizaburo Chubachi +24
In this report, we introduce PLaMo 2, a series of Japanese-focused large language models featuring a hybrid Samba-based architecture that transitions to full attention via continua…
cs.CL2025
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
Kentaro Imajo, Masanori Hirano, Shuji Suzuki +1
Evaluating the open-ended text generation of large language models (LLMs) is challenging because of the lack of a clear ground truth and the high cost of human or LLM-based assessm…
cs.CL2024
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
Preferred Elements, :, Kenshin Abe +18
We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK No…