Searching for a source without gradients: how good is infotaxis and how to beat it
arXiv:2112.10861 · doi:10.1098/rspa.2022.0118
Abstract
Infotaxis is a popular search algorithm designed to track a source of odor in a turbulent environment using information provided by odor detections. To exemplify its capabilities, the source-tracking task was framed as a partially observable Markov decision process consisting in finding, as fast as possible, a stationary target hidden in a 2D grid using stochastic partial observations of the target location. Here we provide an extended review of infotaxis, together with a toolkit for devising better strategies. We first characterize the performance of infotaxis in domains from 1D to 4D. Our results show that, while being suboptimal, infotaxis is reliable (the probability of not reaching the source approaches zero), efficient (the mean search time scales as expected for the optimal strategy), and safe (the tail of the distribution of search times decays faster than any power law, though subexponentially). We then present three possible ways of beating infotaxis, all inspired by methods used in artificial intelligence: tree search, heuristic approximation of the value function, and deep reinforcement learning. The latter is able to find, without any prior human knowledge, the (near) optimal strategy. Altogether, our results provide evidence that the margin of improvement of infotaxis toward the optimal strategy gets smaller as the dimensionality increases.
accepted version
References in corpus (2)
Cited by in corpus (9)
- Optimal active particle navigation meets machine learning
- Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark
- Optimal policies for Bayesian olfactory search in turbulent flows
- Many wrong models approach to localize an odor source in turbulence with static sensors
- A critical assessment of reinforcement learning methods for microswimmer navigation in complex flows
- Optimal trajectories for Bayesian olfactory search in turbulent flows: the low information limit and beyond
- Policy heterogeneity improves collective olfactory search in 3-D turbulence
- Proxitaxis: an adaptive search strategy based on proximity and stochastic resetting
- Learning to traverse convective flows at moderate to high Rayleigh numbers