3 papers
cs.CL2026
From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
Niranjan Chebrolu, Kokil Jaidka, Gerard Christopher Yeo
Complex social behaviors, such as empathy and strategic politeness, are widely assumed to resist the directional decomposition that makes activation steering effective for coarse a…
cs.LG2026
Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
Svetlana Churina, Niranjan Chebrolu, Kokil Jaidka
We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior…
cs.CL2025
Conversations: Love Them, Hate Them, Steer Them
Niranjan Chebrolu, Gerard Christopher Yeo, Kokil Jaidka
Large Language Models (LLMs) demonstrate increasing conversational fluency, yet instilling them with nuanced, human-like emotional expression remains a significant challenge. Curre…