2 papers
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.CY2025
The Goldilocks zone of governing technology: Leveraging uncertainty for responsible quantum practices
Miriam Meckel, Philipp Hacker, Lea Steinacker +4
Emerging technologies challenge conventional governance approaches, especially when uncertainty is not a temporary obstacle but a foundational feature as in quantum computing. This…