2 papers
cs.CL2026
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…
cs.CL2025
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
Atharva Kulkarni, Yuan Zhang, Joel Ruben Antony Moniz +5
Hallucinations pose a significant obstacle to the reliability and widespread adoption of language models, yet their accurate measurement remains a persistent challenge. While many…