3 papers
cs.AI2026
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
Anissa Alloula, Federico Licini, Ava Batchkala +1
LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmark…
cs.LG2025
Representation Invariance and Allocation: When Subgroup Balance Matters
Anissa Alloula, Charles Jones, Zuzanna Wakefield-Skorniewska +2
Unequal representation of demographic groups in training data poses challenges to model generalisation across populations. Standard practice assumes that balancing subgroup represe…
cs.LG2025
Subgroups Matter for Robust Bias Mitigation
Anissa Alloula, Charles Jones, Ben Glocker +1
Despite the constant development of new bias mitigation methods for machine learning, no method consistently succeeds, and a fundamental question remains unanswered: when and why d…