Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Granular feedback merits sophisticated aggregation
Anmol Kagrecha, Henrik Marklund, Potsawee Manakul +2
Human feedback is increasingly used across diverse applications like training AI models, developing recommender systems, and measuring public opinion -- with granular feedback ofte…
cs.LG2025
Misalignment from Treating Means as Ends
Henrik Marklund, Alex Infanger, Benjamin Van Roy
Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about…
cs.LG2024
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
Henrik Marklund, Benjamin Van Roy
As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggest…