4 papers
PLaMo 2 Technical Report
Preferred Networks, :, Kaizaburo Chubachi +24
In this report, we introduce PLaMo 2, a series of Japanese-focused large language models featuring a hybrid Samba-based architecture that transitions to full attention via continua…
Experience Replay with Random Reshuffling
Yasuhiro Fujita
Experience replay is a key component in reinforcement learning for stabilizing learning and improving sample efficiency. Its typical implementation samples transitions with replace…
Entropy Controllable Direct Preference Optimization
Motoki Omura, Yasuhiro Fujita, Toshiki Kataoka
In the post-training of large language models (LLMs), Reinforcement Learning from Human Feedback (RLHF) is an effective approach to achieve generation aligned with human preference…
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
Preferred Elements, :, Kenshin Abe +18
We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK No…