3 papers
cs.CL2025
Interaction Dynamics as a Reward Signal for LLMs
Sian Gooding, Edward Grefenstette
The alignment of Large Language Models (LLMs) for multi-turn conversations typically relies on reward signals derived from the content of the text. This approach, however, overlook…
cs.LG2025
Generative Data Refinement: Just Ask for Better Data
Minqi Jiang, João G. M. Araújo, Will Ellsworth +2
For a fixed parameter size, the capabilities of large models are primarily determined by the quality and quantity of its training data. Consequently, training datasets now grow fas…
cs.CL2025
Writing as a testbed for open ended agents
Sian Gooding, Lucia Lopez-Rivilla, Edward Grefenstette
Open-ended tasks are particularly challenging for LLMs due to the vast solution space, demanding both expansive exploration and adaptable strategies, especially when success lacks…