2 papers
cs.LG2026
Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making
Yue Pei, Hongming Zhang, Chao Gao +7
Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especia…
cs.LG2026
Principled Fast and Meta Knowledge Learners for Continual Reinforcement Learning
Ke Sun, Hongming Zhang, Jun Jin +4
Inspired by the human learning and memory system, particularly the interplay between the hippocampus and cerebral cortex, this study proposes a dual-learner framework comprising a…