2 papers
cs.CL2026
Assessing Reliability of BERT-Based Models on Question Answering Tasks
Pooja Yadav, Priyanka Harjule, Basant Agarwal +1
Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applicati…
cs.AI2026
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik +3
Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and…