1 citations · 1 across the 7 of their papers we have counts for
4 papers · 1 filter
Interaction Dynamics as a Reward Signal for LLMs
Sian Gooding, Edward Grefenstette
The alignment of Large Language Models (LLMs) for multi-turn conversations typically relies on reward signals derived from the content of the text. This approach, however, overlook…
Writing as a testbed for open ended agents
Sian Gooding, Lucia Lopez-Rivilla, Edward Grefenstette
Open-ended tasks are particularly challenging for LLMs due to the vast solution space, demanding both expansive exploration and adaptable strategies, especially when success lacks…
The Impact of Preference Agreement in Reinforcement Learning from Human Feedback: A Case Study in Summarization
Sian Gooding, Hassan Mansoor
Reinforcement Learning from Human Feedback (RLHF) can be used to capture complex and nuanced properties of text generation quality. As a result, the task of text summarization has…
Towards Better Evaluation of Instruction-Following: A Case-Study in Summarization
Ondrej Skopek, Rahul Aralikatte, Sian Gooding +1
Despite recent advances, evaluating how well large language models (LLMs) follow user instructions remains an open problem. While evaluation methods of language models have seen a…