3 papers
cs.LG2025
The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
Ali Falahati, Mohammad Mohammadi Amiri, Kate Larson +1
In self-consuming generative models that train on their own outputs, alignment with user preferences becomes a recursive rather than one-time process. We provide the first formal f…
cs.AI2025
Jackpot! Alignment as a Maximal Lottery
Roberto-Rafael Maura-Rivero, Marc Lanctot, Francesco Visin +1
Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuiti…
cs.MA2024
Soft Condorcet Optimization for Ranking of General Agents
Marc Lanctot, Kate Larson, Michael Kaisers +7
Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a po…