1 paper · 1 filter
William Réveillard, Vasileios Saketos, Alexandre Proutiere +1
Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of how to minimize such intervention while…