4 papers
Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes
Hikaru Shindo, Yu Deng, Teng Cao +5
Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior d…
Kintsugi: Learning Policies by Repairing Executable Knowledge Bases
Teng Cao, Yu Deng, Hikaru Shindo +6
Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy kn…
Checkmating One, by Using Many: Combining Mixture of Experts with MCTS to Improve in Chess
Felix Helfenstein, Johannes Czech, Jannis Blüml +2
In games like chess, strategy evolves dramatically across distinct phases - the opening, middlegame, and endgame each demand different forms of reasoning and decision-making. Yet,…
Better Decisions through the Right Causal World Model
Elisabeth Dillies, Quentin Delfosse, Jannis Blüml +3
Reinforcement learning (RL) agents have shown remarkable performances in various environments, where they can discover effective policies directly from sensory inputs. However, the…