1 paper · 1 filter
Christoph Schnabl, Daniel Hugenroth, Bill Marino +1
Benchmarks are important measures to evaluate safety and compliance of AI models at scale. However, they typically do not offer verifiable results and lack confidentiality for mode…