2 papers
cs.AI2026
Fragility of Value under Imperfect Alignment
Winter Cross, Léo Cymbalista, Alfred Harwood +1
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that hu…
cs.AI2025
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…