Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Interaction Dynamics as a Reward Signal for LLMs
Sian Gooding, Edward Grefenstette
The alignment of Large Language Models (LLMs) for multi-turn conversations typically relies on reward signals derived from the content of the text. This approach, however, overlook…
cs.CL2025
Writing as a testbed for open ended agents
Sian Gooding, Lucia Lopez-Rivilla, Edward Grefenstette
Open-ended tasks are particularly challenging for LLMs due to the vast solution space, demanding both expansive exploration and adaptable strategies, especially when success lacks…