2 citations · 2 across the 1 of their papers we have counts for
1 paper
Longxiang Shi, Shijian Li, Longbing Cao +2
Off-policy reinforcement learning with eligibility traces is challenging because of the discrepancy between target policy and behavior policy. One common approach is to measure the…