1 citations · 1 across the 4 of their papers we have counts for
4 papers
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditio…
A Mechanistic View of Authority Hierarchy in LLM Sycophancy
Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answe…
Automatically Finding Rule-Based Neurons in OthelloGPT
Aditya Singh, Zihang Wen, Srujananjali Medicherla +2
OthelloGPT, a transformer trained to predict valid moves in Othello, provides an ideal testbed for interpretability research. The model is complex enough to exhibit rich computatio…
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise
Shuai Feng, Wei-Chuang Chan, Srishti Chouhan +4
The integration of large language models (LLMs) into global applications necessitates effective cultural alignment for meaningful and culturally-sensitive interactions. Current LLM…