2 papers
cs.CL2025
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
Krithika Ramesh, Daniel Smolyak, Zihao Zhao +4
We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentiall…
cs.LG2024
Improving Equity in Health Modeling with GPT4-Turbo Generated Synthetic Data: A Comparative Study
Daniel Smolyak, Arshana Welivita, Margrét V. Bjarnadóttir +1
Objective. Demographic groups are often represented at different rates in medical datasets. These differences can create bias in machine learning algorithms, with higher levels of…