2 papers
cs.CL2026
ASSERT: A Measurement Pipeline for GenAI Audits
Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…
cs.CL2026
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
Yilun Hua, Giuseppe Castellucci, Peter Schulam +2
Retrieval Augmented Generation (RAG)'s success depends on the utility the LLM derives from the content used for grounding. Quantifying content utility does not have a definitive sp…