5 papers
emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity
Cantao Su, Menan Velayuthan, Esther Ploeger +2
There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…
STEB: Style Text Embedding Benchmark
Rafael Rivera Soto, Anna Wegmann, Cristina Aggazzotti
While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, with each work relying on their o…
Tokenization is Sensitive to Language Variation
Anna Wegmann, Dong Nguyen, David Jurgens
Variation in language is ubiquitous and often systematically linked to regional, social, and contextual factors. Tokenizers split texts into smaller units and might behave differen…
Neurobiber: Fast and Interpretable Stylistic Feature Extraction
Kenan Alkiek, Anna Wegmann, Jian Zhu +1
Linguistic style is pivotal for understanding how texts convey meaning and fulfill communicative purposes, yet extracting detailed stylistic features at scale remains challenging.…
What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
Anna Wegmann, Tijs van den Broek, Dong Nguyen
Best practices for high conflict conversations like counseling or customer support almost always include recommendations to paraphrase the previous speaker. Although paraphrase cla…