3 papers
cs.CL2025
TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
Felipe Nuti, Tim Franzmeyer, João Henriques
Past work has studied the effects of fine-tuning on large language models' (LLMs) overall performance on certain tasks. However, a quantitative and systematic method for analyzing…
cs.LG2024
HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits
Tim Franzmeyer, Aleksandar Shtedritski, Samuel Albanie +3
Benchmarks have been essential for driving progress in machine learning. A better understanding of LLM capabilities on real world tasks is vital for safe development. Designing ade…
cs.LG2024
Select to Perfect: Imitating desired behavior from large multi-agent data
Tim Franzmeyer, Edith Elkind, Philip Torr +2
AI agents are commonly trained with large datasets of demonstrations of human behavior. However, not all behaviors are equally safe or desirable. Desired characteristics for an AI…