6 papers
Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings
Pranav Bhandari, Nicolas Fay, Amitava Datta +2
Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using pref…
Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
Mehwish Nasim, Sanjeevan Selvaganapathy, Neel Ganapathi Sabhahit +6
Many benchmarks show that large language models can answer direct questions about culture. We study a different question: do they also change how they speak when culture is only im…
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
Pranav Bhandari, Nicolas Fay, Sanjeevan Selvaganapathy +3
Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The ne…
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
Pranav Bhandari, Usman Naseem, Mehwish Nasim
Personality steering in large language models (LLMs) commonly relies on injecting trait-specific steering vectors, implicitly assuming that personality traits can be controlled ind…
Can LLM Agents Maintain a Persona in Discourse?
Pranav Bhandari, Nicolas Fay, Michael Wise +4
Large Language Models (LLMs) are widely used as conversational agents, exploiting their capabilities in various sectors such as education, law, medicine, and more. However, LLMs ar…
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
Pranav Bhandari, Usman Naseem, Amitava Datta +2
Psychological assessment tools have long helped humans understand behavioural patterns. While Large Language Models (LLMs) can generate content comparable to that of humans, we exp…