4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.IR2026
Evaluating RAG for French immigration law: a benchmark and baseline study
Annia Abtout, Julien Delaunay, Monika Ewa Rakoczy
International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available benchmark and first comp…
cs.CY2025★ 4 cited
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…