2 papers
cs.LG2024
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
Yunseon Choi, Sangmin Bae, Seonghyun Ban +6
With the advent of foundation models, prompt tuning has positioned itself as an important technique for directing model behaviors and eliciting desired responses. Prompt tuning reg…
cs.LG2024
Mildly Constrained Evaluation Policy for Offline Reinforcement Learning
Linjie Xu, Zhengyao Jiang, Jinyu Wang +2
Offline reinforcement learning (RL) methodologies enforce constraints on the policy to adhere closely to the behavior policy, thereby stabilizing value learning and mitigating the…