3 papers
cs.CL2025
How do we measure privacy in text? A survey of text anonymization metrics
Yaxuan Ren, Krithika Ramesh, Yaxing Yao +1
In this work, we aim to clarify and reconcile metrics for evaluating privacy protection in text through a systematic survey. Although text anonymization is essential for enabling N…
cs.CL2025
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
Krithika Ramesh, Daniel Smolyak, Zihao Zhao +4
We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentiall…
cs.CL2024
Evaluating Differentially Private Synthetic Data Generation in High-Stakes Domains
Krithika Ramesh, Nupoor Gandhi, Pulkit Madaan +3
The difficulty of anonymizing text data hinders the development and deployment of NLP in high-stakes domains that involve private data, such as healthcare and social services. Poor…