Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Fragility of Value under Imperfect Alignment
Winter Cross, Léo Cymbalista, Alfred Harwood +1
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that hu…
cs.AI2026
Calculating Mutual Information between a Reward Maximizer and its Environment
Alfred Harwood, Jose Faustino, Alex Altair
An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of infor…
cs.AI2025
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…