9 papers · 1 filter
Capturing Individual Human Preferences with Reward Features
André Barreto, Vincent Dumoulin, Yiran Mao +6
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…
Plasticity as the Mirror of Empowerment
David Abel, Michael Bowling, André Barreto +13
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has se…
Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
Anthony GX-Chen, Dongyan Lin, Mandana Samiei +4
Language model (LM) agents are increasingly used as autonomous decision-makers which need to actively gather information to guide their decisions. A crucial cognitive skill for suc…
Rejecting Hallucinated State Targets during Planning
Mingde Zhao, Tristan Sylvain, Romain Laroche +2
In planning processes of computational decision-making agents, generative or predictive models are often used as "generators" to propose "targets" representing sets of expected or…
Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning
Martin Klissarov, Akhil Bagaria, Ziyan Luo +3
Developing agents capable of exploring, planning and learning in complex open-ended environments is a grand challenge in artificial intelligence (AI). Hierarchical reinforcement le…
SCAR: Shapley Credit Assignment for More Efficient RLHF
Meng Cao, Shuyuan Zhang, Xiao-Wen Chang +1
Reinforcement Learning from Human Feedback (RLHF) is a widely used technique for aligning Large Language Models (LLMs) with human preferences, yet it often suffers from sparse rewa…