13 papers
AI Agents and the Future of VIS
Chen Zhu-Tian, Nam Wook Kim, Saeed Boorboor +4
Recent advances in agents (i.e., autonomous, goal-driven AI systems that iteratively observe, act, and learn from their environments) offer a fundamentally different approach from…
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
Aryan Sharma, Cutter Dawes, Shivam Raval
When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residu…
Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering
Narmeen Oozeer, Shivam Raval, Philip Quirke +4
Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as…
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
Alexandra Yost, Shreyans Jain, Shivam Raval +6
Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We prop…
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
Cutter Dawes, Aryan Sharma, Angelos Ioannis Lagos +1
Representing and navigating hierarchy is a fundamental primitive of reasoning. Large language models have demonstrated proficiency in a wide variety of tasks requiring hierarchical…
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
Benjamin Maltbie, Shivam Raval
Large language models exhibit sycophantic tendencies, but whether this behavior varies systematically with perceived user demographics is underexplored. Inspired by intersectionali…