Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2026
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
Linh Le, David Williams-King, Mohamed Amine Merzouk +2
Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerab…