2 papers
cs.LG2026
Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs
Elizaveta Tennant, Benjamin Henke, Anita Keshmirian +5
As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional eval…
cs.CL2026
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7
The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…