2 papers
cs.LG2025
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
Francisco Eiras, Eliott Zemour, Eric Lin +1
Large Language Model (LLM) based judges form the underpinnings of key safety evaluation processes such as offline benchmarking, automated red-teaming, and online guardrailing. This…
cs.CL2024
Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Schrasing Tong, Eliott Zemour, Jessica Lu +2
Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in t…