271 citations · 327 across the 9 of their papers we have counts for
5 papers · 1 filter
A Survey of Reinforcement Learning from Human Feedback
Timo Kaufmann, Paul Weng, Viktor Bengs +1
Reinforcement learning from human feedback (RLHF) is a variant of reinforcement learning (RL) that learns from human feedback instead of relying on an engineered reward function. B…
Generalization in Deep RL for TSP Problems via Equivariance and Local Search
Wenbin Ouyang, Yisen Wang, Paul Weng +1
Deep reinforcement learning (RL) has proved to be a competitive heuristic for solving small-sized instances of traveling salesman problems (TSP), but its performance on larger-size…
Fairness in Reinforcement Learning
Paul Weng
Decision support systems (e.g., for ecological conservation) and autonomous systems (e.g., adaptive controllers in smart cities) start to be deployed in real applications. Although…
Exploiting the Sign of the Advantage Function to Learn Deterministic Policies in Continuous Domains
Matthieu Zimmer, Paul Weng
In the context of learning deterministic policies in continuous domains, we revisit an approach, which was first proposed in Continuous Actor Critic Learning Automaton (CACLA) and…
Multi-objective Bandits: Optimizing the Generalized Gini Index
Robert Busa-Fekete, Balazs Szorenyi, Paul Weng +1
We study the multi-armed bandit (MAB) problem where the agent receives a vectorial feedback that encodes many possibly competing objectives to be optimized. The goal of the agent i…