On efficient computation in active inference
arXiv:2307.00504 · doi:10.1016/j.eswa.2024.124315
Abstract
Despite being recognized as neurobiologically plausible, active inference faces difficulties when employed to simulate intelligent behaviour in complex environments due to its computational cost and the difficulty of specifying an appropriate target distribution for the agent. This paper introduces two solutions that work in concert to address these limitations. First, we present a novel planning algorithm for finite temporal horizons with drastically lower computational complexity. Second, inspired by Z-learning from control theory literature, we simplify the process of setting an appropriate target distribution for new and existing active inference planning schemes. Our first approach leverages the dynamic programming algorithm, known for its computational efficiency, to minimize the cost function used in planning through the Bellman-optimality principle. Accordingly, our algorithm recursively assesses the expected free energy of actions in the reverse temporal order. This improves computational efficiency by orders of magnitude and allows precise model learning and planning, even under uncertain conditions. Our method simplifies the planning process and shows meaningful behaviour even when specifying only the agent's final goal state. The proposed solutions make defining a target distribution from a goal state straightforward compared to the more complicated task of defining a temporally informed target distribution. The effectiveness of these methods is tested and demonstrated through simulations in standard grid-world tasks. These advances create new opportunities for various applications.
23 pages, 7 figures. Project repo: https://github.com/aswinpaul/dpefe_2023
References in corpus (7)
- A free energy principle for a particular physics
- The free energy principle made simpler but not too simple
- An empirical evaluation of active inference in multi-armed bandits
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Goal-directed Planning and Goal Understanding by Active Inference: Evaluation Through Simulated and Physical Robot Experiments
- Exploration and preference satisfaction trade-off in reward-free learning
- Modelling non-reinforced preferences using selective attention