2 papers
cs.AI2026
Toward a Theory of Value in AI Alignment
Andrew Smart, Shazeda Ahmed, Jackie Kay +3
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech…
cs.CY2026
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
Rafiya Javed, Cassandra Parent, Jackie Kay +9
Hedging and non-affirmation are behaviors exhibited by large language models (LLMs) that limit the clear endorsement of specific statements. While these behaviors are desirable in…