3 papers
cs.CL2025
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
Krithika Ramesh, Daniel Smolyak, Zihao Zhao +4
We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentiall…
stat.AP2025
Maximizing Predictive Performance for Small Subgroups: Functionally Adaptive Interaction Regularization (FAIR)
Daniel Smolyak, Courtney Paulson, Margrét V. Bjarnadóttir
In many healthcare settings, it is both critical to consider fairness when building analytical applications but also uniquely unacceptable to lower model performance for one group…
cs.LG2024
Improving Equity in Health Modeling with GPT4-Turbo Generated Synthetic Data: A Comparative Study
Daniel Smolyak, Arshana Welivita, Margrét V. Bjarnadóttir +1
Objective. Demographic groups are often represented at different rates in medical datasets. These differences can create bias in machine learning algorithms, with higher levels of…