collaborators

5 papers

cs.LG2026

Strangers to Themselves: What Language Models Say About Themselves Is Generic

Phil Blandfort, Urja Pawar

Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model…

cs.AI2026

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

Urja Pawar, Rajitha Ramanayake, Nabeel Kemal +6

LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named…

cs.CL2026

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

Urja Pawar, Rajitha Ramanayake, Owen O'Neill +5

When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no tru…

cs.LG2026

Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs

Phil Blandfort, Tushar Karayil, Alex McKenzie +3

Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped i…

cs.LG2025

Detecting High-Stakes Interactions with Activation Probes

Alex McKenzie, Urja Pawar, Phil Blandfort +4

Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions -- where the te…