5 citations · 6 across the 3 of their papers we have counts for
3 papers
PILAF: Optimal Human Preference Sampling for Reward Modeling
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng +2
As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerge…
Reward Function Design for Crowd Simulation via Reinforcement Learning
Ariel Kwiatkowski, Vicky Kalogeiton, Julien Pettré +1
Crowd simulation is important for video-games design, since it enables to populate virtual worlds with autonomous avatars that navigate in a human-like manner. Reinforcement learni…
UGAE: A Novel Approach to Non-exponential Discounting
Ariel Kwiatkowski, Vicky Kalogeiton, Julien Pettré +1
The discounting mechanism in Reinforcement Learning determines the relative importance of future and present rewards. While exponential discounting is widely used in practice, non-…