3 papers
cs.CL2026
A Mechanistic View of Authority Hierarchy in LLM Sycophancy
Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answe…
cs.LG2025
Automatically Finding Rule-Based Neurons in OthelloGPT
Aditya Singh, Zihang Wen, Srujananjali Medicherla +2
OthelloGPT, a transformer trained to predict valid moves in Othello, provides an ideal testbed for interpretability research. The model is complex enough to exhibit rich computatio…
cs.AI2025
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise
Shuai Feng, Wei-Chuang Chan, Srishti Chouhan +4
The integration of large language models (LLMs) into global applications necessitates effective cultural alignment for meaningful and culturally-sensitive interactions. Current LLM…