2 papers
cs.CL2026
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
Dhananjay Ashok, Ruth-Ann Armstrong, Jonathan May
Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patte…
cs.CL2026
Textual Entailment is not a Better Bias Metric than Token Probability
Virginia K. Felkner, Allison Lim, Jonathan May
Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-wor…