1 paper · 1 filter
Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer +2
Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI workflows has stayed behind. This…