Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Generating Fair Consensus Statements with Social Choice on Token-Level MDPs
Carter Blair, Kate Larson
Current frameworks for consensus statement generation with large language models lack the inherent structure needed to provide provable fairness guarantees when aggregating diverse…
cs.AI2025
Reflective Verbal Reward Design for Pluralistic Alignment
Carter Blair, Kate Larson, Edith Law
AI agents are commonly aligned with "human values" through reinforcement learning from human feedback (RLHF), where a single reward model is learned from aggregated human feedback…
cs.AI2024
Democratizing Reward Design for Personal and Representative Value-Alignment
Carter Blair, Kate Larson, Edith Law
Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in…