4 papers
STEB: Style Text Embedding Benchmark
Rafael Rivera Soto, Anna Wegmann, Cristina Aggazzotti
While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, with each work relying on their o…
Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion
Rafael Rivera Soto, Barry Chen, Nicholas Andrews
The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated influence operations, motivating the need f…
Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
Rafael Rivera Soto, Barry Chen, Nicholas Andrews
Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the p…
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
Rafael Rivera Soto, Barry Chen, Nicholas Andrews
High-quality paraphrases are easy to produce using instruction-tuned language models or specialized paraphrasing models. Although this capability has a variety of benign applicatio…