2 papers
cs.LG2026
Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes
Hikaru Shindo, Yu Deng, Teng Cao +5
Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior d…
cs.LG2026
Kintsugi: Learning Policies by Repairing Executable Knowledge Bases
Teng Cao, Yu Deng, Hikaru Shindo +6
Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy kn…