83 citations · 105 across the 11 of their papers we have counts for
3 papers · 1 filter
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek +1
The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly tru…
SPHERE: An Evaluation Card for Human-AI Systems
Qianou Ma, Dora Zhao, Xinran Zhao +6
In the era of Large Language Models (LLMs), establishing effective evaluation methods and standards for diverse human-AI interaction systems is increasingly challenging. To encoura…
User-Driven Research of Medical Note Generation Software
Tom Knoll, Francesco Moramarco, Alex Papadopoulos Korfiatis +7
A growing body of work uses Natural Language Processing (NLP) methods to automatically generate medical notes from audio recordings of doctor-patient consultations. However, there…