2 papers
cs.CL2025
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
Jennifer Healey, Laurie Byrum, Md Nadeem Akhtar +2
LLM evaluation is challenging even the case of base models. In real world deployments, evaluation is further complicated by the interplay of task specific prompts and experiential…
cs.CL2024
Evaluating Nuanced Bias in Large Language Model Free Response Answers
Jennifer Healey, Laurie Byrum, Md Nadeem Akhtar +1
Pre-trained large language models (LLMs) can now be easily adapted for specific business purposes using custom prompts or fine tuning. These customizations are often iteratively re…