1 citations · 1 across the 8 of their papers we have counts for
9 papers
Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation
Gijs Kassenaar, Zhao Yang, Vincent François-Lavet
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computatio…
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
Taewoon Kim, Vincent François-Lavet, Michael Cochez
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic m…
Reinforcement Learning from Cross-domain Videos with Video Prediction Model
Zhao Yang, Xinrui Zu, Jacob E. Kooi +5
Reinforcement learning from expert videos across visually distinct domains is challenging due to the absence of reward signals and the presence of domain gaps. We introduce XIPER (…
Guiding Skill Discovery with Foundation Models
Zhao Yang, Thomas M. Moerland, Mike Preuss +3
Learning diverse skills without hand-crafted reward functions could accelerate reinforcement learning in downstream tasks. However, existing skill discovery methods focus solely on…
Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning
Alessio Arcudi, Davide Sartor, Alberto Sinigaglia +2
This paper introduces MANGO (Multilayer Abstraction for Nested Generation of Options), a novel hierarchical reinforcement learning framework designed to address the challenges of l…
Novel RL approach for efficient Elevator Group Control Systems
Nathan Vaartjes, Vincent Francois-Lavet
Efficient elevator traffic management in large buildings is critical for minimizing passenger travel times and energy consumption. Because heuristic- or pattern-detection-based con…