1 paper
Buyun Liang, Jinqi Luo, Liangzu Peng +6
Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, making it important to systematically evaluate their reliability…