Publications (12)
Subjective Code Preferences in Experts and Large Language Models
Anna Mokhova, Subhabrata Dutta, Iryna Gurevych +1
Large Language Models (LLMs) have become increasingly popular for coding tasks, with subjective coding preferences being an essential element to adapt to programmers' personal need…
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
PatrÃcia Schmidtová, Saad Mahamood, Simone Balloccu +6
Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…
PhilHumans: Benchmarking Machine Learning for Personal Health
Vadim Liventsev, Vivek Kumar, Allmin Pradhap Singh Susaiyah +14
The use of machine learning in Healthcare has the potential to improve patient outcomes as well as broaden the reach and affordability of Healthcare. The history of other applicati…
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
ZdenÄk Kasner, Vilém Zouhar, PatrÃcia Schmidtová +7
Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek +1
The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly tru…
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
Simone Balloccu, PatrÃcia Schmidtová, Mateusz Lango +1
Natural Language Processing (NLP) research is increasingly focusing on the use of Large Language Models (LLMs), with some of the most popular ones being either fully or partially c…
Ask the experts: sourcing high-quality datasets for nutritional counselling through Human-AI collaboration
Simone Balloccu, Ehud Reiter, Vivek Kumar +2
Large Language Models (LLMs), with their flexible generation abilities, can be powerful data sources in domains with few or no available corpora. However, problems like hallucinati…
Comparing informativeness of an NLG chatbot vs graphical app in diet-information domain
Simone Balloccu, Ehud Reiter
Visual representation of data like charts and tables can be challenging to understand for readers. Previous work showed that combining visualisations with text can improve the comm…
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
Federico Tiblias, Irina Bigoulaeva, Jingcheng Niu +2
The linear representation hypothesis states that language models (LMs) encode concepts as directions in their latent space, forming organized, multidimensional manifolds. Prior wor…
factgenie: A Framework for Span-based Evaluation of Generated Texts
ZdenÄk Kasner, OndÅej Plátek, PatrÃcia Schmidtová +2
We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…
The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation
Doan Nam Long Vu, Simone Balloccu
Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evaluate 12 open-weight vision-language models…
How are you? Introducing stress-based text tailoring
Simone Balloccu, Ehud Reiter, Alexandra Johnstone +1
Can stress affect not only your life but also how you read and interpret a text? Healthcare has shown evidence of such dynamics and in this short paper we discuss customising texts…