Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Representation over Routing: Diagnosing Temporal Routing Pathologies in Multi-Timescale PPO
Jing Sun
Temporal credit assignment in reinforcement learning is often approached by introducing value estimates at multiple discount factors. A natural next step is to let the actor dynami…
cs.LG2024
Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
Songjun Tu, Jingbo Sun, Qichao Zhang +2
Preference-based reinforcement learning (PbRL) provides a powerful paradigm to avoid meticulous reward engineering by learning rewards based on human preferences. However, real-tim…