2 papers
cs.CL2026
Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
Isotta Magistrali, Chen Shani
The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's dec…
cs.LG2026
Subliminal Signals in Preference Labels
Isotta Magistrali, Frédéric Berdoz, Sam Dauncey +1
As AI systems approach superhuman capabilities, scalable oversight increasingly relies on LLM-as-a-judge frameworks where models evaluate and guide each other's training. A core as…