2 papers
cs.LG2026
ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
Zijun Xie, Yuyang You, Yongzhi Li +8
Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their…
cs.RO2025
DTCCL: Disengagement-Triggered Contrastive Continual Learning for Autonomous Bus Planners
Yanding Yang, Weitao Zhou, Jinhai Wang +8
Autonomous buses run on fixed routes but must operate in open, dynamic urban environments. Disengagement events on these routes are often geographically concentrated and typically…