1 paper · 1 filter
Guang Yang, Xinyang Liu
Large Language Models (LLMs) have shown remarkable progress in multiple-choice question answering (MCQA), but their inherent unreliability, such as hallucination and overconfidence…