6 citations · 10 across the 5 of their papers we have counts for
5 papers · 1 filter
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifica…
Understanding the Effectiveness of Very Large Language Models on Dialog Evaluation
Jessica Huynh, Cathy Jiao, Prakhar Gupta +4
Language models have steadily increased in size over the past few years. They achieve a high level of performance on various natural language processing (NLP) tasks such as questio…
DialCrowd 2.0: A Quality-Focused Dialog System Crowdsourcing Toolkit
Jessica Huynh, Ting-Rui Chiang, Jeffrey Bigham +1
Dialog system developers need high-quality data to train, fine-tune and assess their systems. They often use crowdsourcing for this since it provides large quantities of data from…
A Survey of NLP-Related Crowdsourcing HITs: what works and what does not
Jessica Huynh, Jeffrey Bigham, Maxine Eskenazi
Crowdsourcing requesters on Amazon Mechanical Turk (AMT) have raised questions about the reliability of the workers. The AMT workforce is very diverse and it is not possible to mak…
SAPPHIRE: Approaches for Enhanced Concept-to-Text Generation
Steven Y. Feng, Jessica Huynh, Chaitanya Narisetty +2
We motivate and propose a suite of simple but effective improvements for concept-to-text generation called SAPPHIRE: Set Augmentation and Post-hoc PHrase Infilling and REcombinatio…