Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Adaptable Hindsight Experience Replay for Search-Based Learning
Alexandros Vazaios, Jannis Brugger, Cedric Derstroff +2
AlphaZero-like Monte Carlo Tree Search systems, originally introduced for two-player games, dynamically balance exploration and exploitation using neural network guidance. This com…
cs.LG2025
Deep Reinforcement Learning via Object-Centric Attention
Jannis Blüml, Cedric Derstroff, Bjarne Gregori +3
Deep reinforcement learning agents, trained on raw pixel inputs, often fail to generalize beyond their training environments, relying on spurious correlations and irrelevant backgr…
cs.LG2025
Polynomial Regret Concentration of UCB for Non-Deterministic State Transitions
Can Cömer, Jannis Blüml, Cedric Derstroff +1
Monte Carlo Tree Search (MCTS) has proven effective in solving decision-making problems in perfect information settings. However, its application to stochastic and imperfect inform…