activity
20232025
most citedThis is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models

3 citations · 3 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL2025

Challenging the Abilities of Large Language Models in Italian: a Community Initiative

Malvina Nissim, Danilo Croce, Viviana Patti +78

The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of t…

cs.CL2025

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

Jeremy Barnes, Naiara Perez, Alba Bonet-Jover +1

Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these…

cs.CL2024

NoticIA: A Clickbait Article Summarization Dataset in Spanish

Iker García-Ferrero, Begoña Altuna

We present NoticIA, a dataset consisting of 850 Spanish news articles featuring prominent clickbait headlines, each paired with high-quality, single-sentence generative summarizati…

cs.CL2024

A Hard Nut to Crack: Idiom Detection with Conversational Large Language Models

Francesca De Luca Fornaciari, Begoña Altuna, Itziar Gonzalez-Dios +1

In this work, we explore idiomatic language processing with Large Language Models (LLMs). We introduce the Idiomatic language Test Suite IdioTS, a new dataset of difficult examples…

cs.CL20233 cited

This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models

Iker García-Ferrero, Begoña Altuna, Javier Álvez +2

Although large language models (LLMs) have apparently acquired a certain level of grammatical knowledge and the ability to make generalizations, they fail to interpret negation, a…