Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
Dhananjay Ashok, Ruth-Ann Armstrong, Jonathan May
Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patte…
cs.CL2026
Textual Entailment is not a Better Bias Metric than Token Probability
Virginia K. Felkner, Allison Lim, Jonathan May
Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-wor…
cs.CL2024
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
Virginia K. Felkner, Jennifer A. Thompson, Jonathan May
Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work h…