11 papers
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
Donghu Kim, Youngdo Lee, Hojoon Lee +6
Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challeng…
Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions
Bingxu Liu, Jiashun Liu, Johan Obando-Ceron +5
While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stat…
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
Ali Saheb Pasand, Johan Obando-Ceron, Aaron Courville +2
Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data distributions evolve over time. We show…
Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
Johan Obando-Ceron, Walter Mayor, Samuel Lavoie +3
Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment parallelization; unfortunately, these can someti…
A Mechanistic Analysis of Looped Reasoning Language Models
Hugh Blayney, Ãlvaro Arroyo, Johan Obando-Ceron +4
Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM's layers in the latent di…
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
Jiashun Liu, Zihao Wu, Johan Obando-Ceron +3
Deep reinforcement learning (RL) agents frequently suffer from neuronal activity loss, which impairs their ability to adapt to new data and learn continually. A common method to qu…