1 paper · 1 filter
Nandiraju Gireesh, Yuanliang Ju, He Wang
Offline-to-online reinforcement learning with action chunking eliminates multi-step off-policy bias and enables temporally coherent exploration, but all existing methods use a fixe…