8 papers
Policy Gradient for Continuous-Time Robust Markov Decision Processes
Tanya Veeravalli, David M. Bossens, Atsushi Nitanda
The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under worst-case transition dynamic…
LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions
Xin He, Junxi Shen, Yuchen Mou +4
Classical models of opinion dynamics assume human participants with bounded rationality and limited coordination. The rise of LLM-based agents introduces a qualitative shift: agent…
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
David M. Bossens, Atsushi Nitanda
Safety is an essential requirement for reinforcement learning systems. The newly emerging framework of robust constrained Markov decision processes allows learning policies that sa…
The Digital Ecosystem of Beliefs: does evolution favour AI over humans?
David M. Bossens, Shanshan Feng, Yew-Soon Ong
As AI systems are integrated into social networks, there are AI safety concerns that AI-generated content may dominate the web, e.g. in popularity or impact on beliefs. To understa…
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
Zhenglin Wan, Xingrui Yu, David Mark Bossens +5
Imitation learning (IL) has shown promise in various applications (e.g. robot locomotion) but is often limited to learning a single expert policy, constraining behavior diversity a…
Quantum Policy Gradient in Reproducing Kernel Hilbert Space
David M. Bossens, Kishor Bharti, Jayne Thompson
Parametrised quantum circuits offer expressive and data-efficient representations for machine learning. Due to quantum states residing in a high-dimensional Hilbert space, parametr…