Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2025
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
Hadi Nekoei, Alexandre Blondin Massé, Rachid Hassani +2
Reinforcement learning (RL) is a powerful framework for optimizing decision-making in complex systems under uncertainty, an essential challenge in real-world settings, particularly…