activity
20242026
collaborators

5 papers

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.LG2026

To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Learning, Except In Heavy Truncation Scenarios

Sotaro Takeshita, Yurina Takeshita, Simone Paolo Ponzetto +1

Matryoshka Representation Learning (MRL) is a widely adopted approach for training text encoders so they provide useful text representations at various sizes, available by simply t…

cs.LG2025

Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks

Sotaro Takeshita, Yurina Takeshita, Daniel Ruffinelli +1

In this paper, we study the surprising impact that truncating text embeddings has on downstream performance. We consistently observe across 6 state-of-the-art text encoders and 26…

cs.CL2025

Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect

Alina Klerings, Jannik Brinkmann, Daniel Ruffinelli +1

Large language models (LLMs) are able to generate grammatically well-formed text, but how do they encode their syntactic knowledge internally? While prior work has focused largely…

cs.DL2024

Enriching Social Science Research via Survey Item Linking

Tornike Tsereteli, Daniel Ruffinelli, Simone Paolo Ponzetto

Questions within surveys, called survey items, are used in the social sciences to study latent concepts, such as the factors influencing life satisfaction. Instead of using explici…