papers

Publications (12)

cs.HC2026

Subjective Code Preferences in Experts and Large Language Models

Anna Mokhova, Subhabrata Dutta, Iryna Gurevych +1

Large Language Models (LLMs) have become increasingly popular for coding tasks, with subjective coding preferences being an essential element to adapt to programmers' personal need…

cs.CL2024

Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices

Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6

Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…

cs.LG2024

PhilHumans: Benchmarking Machine Learning for Personal Health

Vadim Liventsev, Vivek Kumar, Allmin Pradhap Singh Susaiyah +14

The use of machine learning in Healthcare has the potential to improve patient outcomes as well as broaden the reach and affordability of Healthcare. The history of other applicati…

cs.CL2026

LLMs as Span Annotators: A Comparative Study of LLMs and Humans

Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová +7

Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…

cs.HC2025

When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition

Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek +1

The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly tru…

cs.CL2024

Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Simone Balloccu, Patrícia Schmidtová, Mateusz Lango +1

Natural Language Processing (NLP) research is increasingly focusing on the use of Large Language Models (LLMs), with some of the most popular ones being either fully or partially c…

cs.CL2024

Ask the experts: sourcing high-quality datasets for nutritional counselling through Human-AI collaboration

Simone Balloccu, Ehud Reiter, Vivek Kumar +2

Large Language Models (LLMs), with their flexible generation abilities, can be powerful data sources in domains with few or no available corpora. However, problems like hallucinati…

cs.CL2022

Comparing informativeness of an NLG chatbot vs graphical app in diet-information domain

Simone Balloccu, Ehud Reiter

Visual representation of data like charts and tables can be challenging to understand for readers. Previous work showed that combining visualisations with text can improve the comm…

cs.AI2026

Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling

Federico Tiblias, Irina Bigoulaeva, Jingcheng Niu +2

The linear representation hypothesis states that language models (LMs) encode concepts as directions in their latent space, forming organized, multidimensional manifolds. Prior wor…

cs.CL2024

factgenie: A Framework for Span-based Evaluation of Generated Texts

Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2

We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…

cs.AI2026

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Doan Nam Long Vu, Simone Balloccu

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evaluate 12 open-weight vision-language models…

cs.CL2020

How are you? Introducing stress-based text tailoring

Simone Balloccu, Ehud Reiter, Alexandra Johnstone +1

Can stress affect not only your life but also how you read and interpret a text? Healthcare has shown evidence of such dynamics and in this short paper we discuss customising texts…