1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2024
Dynamic Weight Adjusting Deep Q-Networks for Real-Time Environmental Adaptation
Xinhao Zhang, Jinghan Zhang, Wujun Si +1
Deep Reinforcement Learning has shown excellent performance in generating efficient solutions for complex tasks. However, its efficacy is often limited by static training modes and…
cs.CL2024★ 1 cited
Prototypical Reward Network for Data-Efficient RLHF
Jinghan Zhang, Xiting Wang, Yiqiao Jin +3
The reward model for Reinforcement Learning from Human Feedback (RLHF) has proven effective in fine-tuning Large Language Models (LLMs). Notably, collecting human feedback for RLHF…