collaborators

12 papers

cs.LG2026

Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes

Hikaru Shindo, Yu Deng, Teng Cao +5

Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior d…

cs.CV2026

STORM: Segment, Track, and Object Re-Localization from a Single Image

Yu Deng, Teng Cao, Hikaru Shindo +3

Accurate 6D pose estimation and tracking are core capabilities for physical AI systems, yet real-world deployment remains brittle and labor-intensive. Many pipelines rely on CAD mo…

cs.LG2026

Kintsugi: Learning Policies by Repairing Executable Knowledge Bases

Teng Cao, Yu Deng, Hikaru Shindo +6

Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy kn…

cs.AI2026

GRAIL: Autonomous Concept Grounding for Neuro-Symbolic Reinforcement Learning

Hikaru Shindo, Henri Rößler, Quentin Delfosse +1

Neuro-symbolic Reinforcement Learning (NeSy-RL) combines symbolic reasoning with gradient-based optimization to achieve interpretable and generalizable policies. Relational concept…

cs.LG2026

LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Lukas Helff, Quentin Delfosse, David Steinmann +6

As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode emerges: LLMs gaming verifi…

cs.AI2026

Boosting deep Reinforcement Learning using pretraining with Logical Options

Zihan Ye, Phil Chau, Raban Emunds +5

Deep reinforcement learning agents are often misaligned, as they over-exploit early reward signals. Recently, several symbolic approaches have addressed these challenges by encodin…