1 citations · 1 across the 4 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Boosting Deductive Reasoning with Step Signals In RLHF
Jialian Li, Yipin Zhang, Wei Shen +3
Logical reasoning is a crucial task for Large Language Models (LLMs), enabling them to tackle complex problems. Among reasoning tasks, multi-step reasoning poses a particular chall…
cs.LG2024
Reward-Robust RLHF in LLMs
Yuzi Yan, Xingzhou Lou, Jialian Li +6
As Large Language Models (LLMs) continue to progress toward more advanced forms of intelligence, Reinforcement Learning from Human Feedback (RLHF) is increasingly seen as a key pat…