3 papers
cs.CL2025
Text2Grad: Reinforcement Learning from Natural Language Feedback
Hanyang Wang, Lu Wang, Chaoyun Zhang +5
Traditional RLHF optimizes language models with coarse, scalar rewards that mask the fine-grained reasons behind success or failure, leading to slow and opaque learning. Recent wor…
cs.LG2025
Fine-Tuning without Performance Degradation
Han Wang, Adam White, Martha White
Fine-tuning policies learned offline remains a major challenge in application domains. Monotonic performance improvement during \emph{fine-tuning} is often challenging, as agents t…
cs.LG2025
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
Lingwei Zhu, Han Wang, Yukie Nagai
Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the…