1 citations · 2 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning
Gong Gao, Xiao Lai, Ziqi Xie +3
Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-differ…
cs.LG2026★ 1 cited
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
Gong Gao, Weidong Zhao, Xianhui Liu +1
Existing value-based online reinforcement learning (RL) algorithms suffer from slow policy exploitation due to ineffective exploration and delayed policy updates. To address these…