1 paper
Miguel Moura Ramos, Patrick Fernandes, António Farinhas +1
Reinforcement learning from human feedback (RLHF) is a recent technique to improve the quality of the text generated by a language model, making it closer to what humans would gene…