Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
Georgios Papoudakis, Thomas Coste, Jianye Hao +2
Reinforcement learning (RL) using foundation models for policy approximations in multi-turn tasks remains challenging. We identify two main limitations related to sparse reward set…
cs.LG2024
Bayesian Reward Models for LLM Alignment
Adam X. Yang, Maxime Robeyns, Thomas Coste +4
To ensure that large language model (LLM) responses are helpful and non-toxic, a reward model trained on human preference data is usually used. LLM responses with high rewards are…