2 citations · 3 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 2 cited
Human-Calibrated Automated Testing and Validation of Generative Language Models
Agus Sudjianto, Aijun Zhang, Srinivas Neppalli +2
This paper introduces a comprehensive framework for the evaluation and validation of generative language models (GLMs), with a focus on Retrieval-Augmented Generation (RAG) systems…
cs.CL2024
Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting
Ying Li, Rahul Singh, Tarun Joshi +1
Recent work in behavioral testing for natural language processing (NLP) models, such as Checklist, is inspired by related paradigms in software engineering testing. They allow eval…