4 papers
Target-Aligned Reinforcement Learning
Leonard S. Pleiss, James Harrison, Maximilian Schiffer
Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training. While effective, this mechanism intro…
Disentangling generalization and memorization in large language models using chess
Leonard S. Pleiss, Maximilian Schiffer, Robert K. von Weizsaecker
Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genuine reasoning ability. We introduce ch…
Synthetic Monitoring Environments for Reinforcement Learning
Leonard Pleiss, Carolin Schmidt, Maximilian Schiffer
Reinforcement Learning (RL) lacks benchmarks that enable precise, white-box diagnostics of agent behavior. Current environments often entangle complexity factors and lack ground-tr…
Reliability-Adjusted Prioritized Experience Replay
Leonard S. Pleiss, Tobias Sutter, Maximilian Schiffer
Experience replay enables data-efficient learning from past experiences in online reinforcement learning agents. Traditionally, experiences were sampled uniformly from a replay buf…