5 papers
Blackbird Language Matrices: A Framework to Investigate the Linguistic Competence of Language Models
Paola Merlo, Chunyang Jiang, Giuseppe Samo +1
This article describes a novel language task, the Blackbird Language Matrices (BLM) task, inspired by intelligence tests, and illustrates the BLM datasets, their construction and b…
Challenging the Abilities of Large Language Models in Italian: a Community Initiative
Malvina Nissim, Danilo Croce, Viviana Patti +78
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of t…
Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
Vivi Nastase, Paola Merlo
Transformer models learn to encode and decode an input text, and produce contextual token embeddings as a side-effect. The mapping from language into the embedding space maps words…
Exploring Italian sentence embeddings properties through multi-tasking
Vivi Nastase, Giuseppe Samo, Chunyang Jiang +1
We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several…
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
Vivi Nastase, Chunyang Jiang, Giuseppe Samo +1
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the…