3 papers
cs.LG2025
Experience Replay with Random Reshuffling
Yasuhiro Fujita
Experience replay is a key component in reinforcement learning for stabilizing learning and improving sample efficiency. Its typical implementation samples transitions with replace…
cs.CL2025
PLaMo 2 Technical Report
Preferred Networks, :, Kaizaburo Chubachi +24
In this report, we introduce PLaMo 2, a series of Japanese-focused large language models featuring a hybrid Samba-based architecture that transitions to full attention via continua…
cs.LG2025
Entropy Controllable Direct Preference Optimization
Motoki Omura, Yasuhiro Fujita, Toshiki Kataoka
In the post-training of large language models (LLMs), Reinforcement Learning from Human Feedback (RLHF) is an effective approach to achieve generation aligned with human preference…