2 papers
cs.LG2026
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Jonathan Colaço Carr, Prakash Panangaden +2
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
cs.LG2025
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Samuel Garcin, Trevor McInroe, Pablo Samuel Castro +4
Extracting relevant information from a stream of high-dimensional observations is a central challenge for deep reinforcement learning agents. Actor-critic algorithms add further co…