5 papers
Strangers to Themselves: What Language Models Say About Themselves Is Generic
Phil Blandfort, Urja Pawar
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model…
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
Urja Pawar, Rajitha Ramanayake, Nabeel Kemal +6
LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named…
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
Urja Pawar, Rajitha Ramanayake, Owen O'Neill +5
When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no tru…
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
Phil Blandfort, Tushar Karayil, Alex McKenzie +3
Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped i…
Detecting High-Stakes Interactions with Activation Probes
Alex McKenzie, Urja Pawar, Phil Blandfort +4
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions -- where the te…