Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2025
Causal Cartographer: From Mapping to Reasoning Over Counterfactual Worlds
Gaël Gendron, Jože M. Rožanec, Michael Witbrock +1
Causal world models are systems that can answer counterfactual questions about an environment of interest, i.e. predict how it would have evolved if an arbitrary subset of events h…