6 papers
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
Yang Chen, Menglin Zou, Jiaqi Zhang +6
Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates b…
Causal Cartographer: From Mapping to Reasoning Over Counterfactual Worlds
Gaël Gendron, Jože M. Rožanec, Michael Witbrock +1
Causal world models are systems that can answer counterfactual questions about an environment of interest, i.e. predict how it would have evolved if an arbitrary subset of events h…
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
Qiming Bao, Gael Gendron, Alex Yuxuan Peng +5
Large language models (LLMs), such as LLaMA, Alpaca, Vicuna, GPT-3.5 and GPT-4, have advanced the performance of AI systems on various natural language processing tasks to human-li…
Counterfactual Causal Inference in Natural Language with Large Language Models
Gaël Gendron, Jože M. Rožanec, Michael Witbrock +1
Causal structure discovery methods are commonly applied to structured data where the causal variables are known and where statistical testing can be used to assess the causal relat…
Robust Domain Generalisation with Causal Invariant Bayesian Neural Networks
Gaël Gendron, Michael Witbrock, Gillian Dobbie
Deep neural networks can obtain impressive performance on various tasks under the assumption that their training domain is identical to their target domain. Performance can drop dr…