1 citations · 1 across the 13 of their papers we have counts for
5 papers · 1 filter
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
Benjamin Maltbie, Shivam Raval
Large language models exhibit sycophantic tendencies, but whether this behavior varies systematically with perceived user demographics is underexplored. Inspired by intersectionali…
Curveball Steering: The Right Direction To Steer Isn't Always Linear
Shivam Raval, Hae Jin Song, Linlin Wu +4
Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the L…
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
Idhant Gulati, Shivam Raval
Lifelong multimodal agents must continuously adapt to new tasks through post-training, but this creates a fundamental tension between acquiring capabilities and preserving safety a…
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
Alexandra Yost, Shreyans Jain, Shivam Raval +6
Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We prop…
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
Gerard Boxo, Aman Neelappa, Shivam Raval
White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting tex…