3 papers
cs.LG2025
A Convolution and Attention Based Encoder for Reinforcement Learning under Partial Observability
Wuhao Wang, Zhiyong Chen
Partially Observable Markov Decision Processes (POMDPs) remain a core challenge in reinforcement learning due to incomplete state information. We address this by reformulating POMD…
cs.AI2025
Success in Humanoid Reinforcement Learning under Partial Observation
Wuhao Wang, Zhiyong Chen
Reinforcement learning has been widely applied to robotic control, but effective policy learning under partial observability remains a major challenge, especially in high-dimension…
cs.LG2024
Multi-State TD Target for Model-Free Reinforcement Learning
Wuhao Wang, Zhiyong Chen, Lepeng Zhang
Temporal difference (TD) learning is a fundamental technique in reinforcement learning that updates value estimates for states or state-action pairs using a TD target. This target…