Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Aligning AI Agents via Information-Directed Sampling
Hong Jun Jeon, Benjamin Van Roy
The staggering feats of AI systems have brought to attention the topic of AI Alignment: aligning a "superintelligent" AI agent's actions with humanity's interests. Many existing fr…
cs.LG2024
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
Henrik Marklund, Benjamin Van Roy
As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggest…