2 papers
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.CY2026
'AI Alignment' Encompasses Competing Technical Priorities
Tushita Jha, Rory Svarc, Mateusz BagiÅski
The ML literature contains many distinct concepts falling under the heading of 'AI alignment'. After noting three concepts of AI alignment in the context of their corresponding res…