2 papers
cs.LG2025
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Sheng Ouyang, Yulan Hu, Ge Chen +3
Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exh…
cs.AI2024
GUNDAM: Aligning Large Language Models with Graph Understanding
Sheng Ouyang, Yulan Hu, Ge Chen +1
Large Language Models (LLMs) have achieved impressive results in processing text data, which has sparked interest in applying these models beyond textual data, such as graphs. In t…