3 papers
cs.AI2026
Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification
Qian Ma, Anna Squicciarini, Sarah Rajtmajer
Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation met…
cs.CR2026
Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
Qian Ma, Sarah Rajtmajer
Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which…
cs.CL2025
The study of short texts in digital politics: Document aggregation for topic modeling
Nitheesha Nakka, Omer F. Yalcin, Bruce A. Desmarais +2
Statistical topic modeling is widely used in political science to study text. Researchers examine documents of varying lengths, from tweets to speeches. There is ongoing debate on…