1 paper
Sarra Gharsallah, Adele Robaldo, Mariia Tokareva +5
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuali…