1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Katarzyna Kobalczyk, Mihaela van der Schaar
Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs.…