1 paper · 1 filter
Chen Feng, Minghe Shen, Ananth Balashankar +2
Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalabil…