Parisa Rabbani, Priyam Sahoo, Ruben Mathew +4
LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that LLMs judge identical claims dif…