2 papers
cs.LG2026
Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making
Yue Pei, Hongming Zhang, Chao Gao +7
Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especia…
cs.LG2025
-DQN: Improving Deep Q-Learning By Evolving the Behavior
Hongming Zhang, Fengshuo Bai, Chenjun Xiao +3
While many sophisticated exploration methods have been proposed, their lack of generality and high computational cost often lead researchers to favor simpler methods like -gree…