5 papers
Subjective Code Preferences in Experts and Large Language Models
Anna Mokhova, Subhabrata Dutta, Iryna Gurevych +1
Large Language Models (LLMs) have become increasingly popular for coding tasks, with subjective coding preferences being an essential element to adapt to programmers' personal need…
Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions
Doan Nam Long Vu, Simone Balloccu
Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging c…
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová +7
Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6
Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…
factgenie: A Framework for Span-based Evaluation of Generated Texts
Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2
We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…