6 papers
Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
Giuseppe Samo, Paola Merlo
This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Itali…
Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
Giuseppe Samo, Paola Merlo
Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns…
Blackbird Language Matrices: A Framework to Investigate the Linguistic Competence of Language Models
Paola Merlo, Chunyang Jiang, Giuseppe Samo +1
This article describes a novel language task, the Blackbird Language Matrices (BLM) task, inspired by intelligence tests, and illustrates the BLM datasets, their construction and b…
Modelling the Morphology of Verbal Paradigms: A Case Study in the Tokenization of Turkish and Hebrew
Giuseppe Samo, Paola Merlo
We investigate how transformer models represent complex verb paradigms in Turkish and Modern Hebrew, concentrating on how tokenization strategies shape this ability. Using the Blac…
Exploring Italian sentence embeddings properties through multi-tasking
Vivi Nastase, Giuseppe Samo, Chunyang Jiang +1
We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several…
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
Vivi Nastase, Chunyang Jiang, Giuseppe Samo +1
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the…