3 citations · 3 across the 19 of their papers we have counts for
11 papers · 1 filter
Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes
Hikaru Shindo, Yu Deng, Teng Cao +5
Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior d…
Kintsugi: Learning Policies by Repairing Executable Knowledge Bases
Teng Cao, Yu Deng, Hikaru Shindo +6
Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy kn…
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Lukas Helff, Quentin Delfosse, David Steinmann +6
As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode emerges: LLMs gaming verifi…
Deep Reinforcement Learning Agents are not even close to Human Intelligence
Quentin Delfosse, Jannis Blüml, Fabian Tatai +6
Deep reinforcement learning (RL) agents achieve impressive results in a wide variety of tasks, but they lack zero-shot adaptation capabilities. While most robustness evaluations fo…
Deep Reinforcement Learning via Object-Centric Attention
Jannis Blüml, Cedric Derstroff, Bjarne Gregori +3
Deep reinforcement learning agents, trained on raw pixel inputs, often fail to generalize beyond their training environments, relying on spurious correlations and irrelevant backgr…
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
Hector Kohler, Quentin Delfosse, Waris Radji +2
There exist applications of reinforcement learning like medicine where policies need to be ''interpretable'' by humans. User studies have shown that some policy classes might be mo…