4 papers
Improving Cross-Lingual Token Representations by Adding a Pinch of SALT
Guillem Ramírez
Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource setti…
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot +17
Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSO…
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
Guillem Ramírez, Alexandra Birch, Ivan Titov
Large language models (LLMs) are primarily accessed via commercial APIs, but this often requires users to expose their data to service providers. In this paper, we explore how user…
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
Guillem Ramírez, Alexandra Birch, Ivan Titov
Researchers and practitioners operating on a limited budget face the cost-performance trade-off dilemma. The challenging decision often centers on whether to use a large LLM with b…