activity
20242026
collaborators

5 papers

cs.CL2026

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Cantao Su, Menan Velayuthan, Esther Ploeger +2

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…

cs.CL2026

STEB: Style Text Embedding Benchmark

Rafael Rivera Soto, Anna Wegmann, Cristina Aggazzotti

While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, with each work relying on their o…

cs.CL2025

Tokenization is Sensitive to Language Variation

Anna Wegmann, Dong Nguyen, David Jurgens

Variation in language is ubiquitous and often systematically linked to regional, social, and contextual factors. Tokenizers split texts into smaller units and might behave differen…

cs.CL2025

Neurobiber: Fast and Interpretable Stylistic Feature Extraction

Kenan Alkiek, Anna Wegmann, Jian Zhu +1

Linguistic style is pivotal for understanding how texts convey meaning and fulfill communicative purposes, yet extracting detailed stylistic features at scale remains challenging.…

cs.CL2024

What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs

Anna Wegmann, Tijs van den Broek, Dong Nguyen

Best practices for high conflict conversations like counseling or customer support almost always include recommendations to paraphrase the previous speaker. Although paraphrase cla…