4 papers
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
Omnilingual SONAR Team, João Maria Janeiro, Pere-LluÃs Huguet Cabot +17
Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSO…
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
Guillem RamÃrez, Alexandra Birch, Ivan Titov
Large language models (LLMs) are primarily accessed via commercial APIs, but this often requires users to expose their data to service providers. In this paper, we explore how user…
On a Novel Application of Wasserstein-Procrustes for Unsupervised Cross-Lingual Learning
Guillem RamÃrez, Rumen Dangovski, Preslav Nakov +1
The emergence of unsupervised word embeddings, pre-trained on very large monolingual text corpora, is at the core of the ongoing neural revolution in Natural Language Processing (N…
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
Guillem RamÃrez, Alexandra Birch, Ivan Titov
Researchers and practitioners operating on a limited budget face the cost-performance trade-off dilemma. The challenging decision often centers on whether to use a large LLM with b…