6 citations · 8 across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
Yash Aggarwal, Atmika Gorti, Vinija Jain +3
Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "…
cs.LG2025
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
Deven Mahesh Mistry, Anooshka Bajaj, Yash Aggarwal +2
We investigate in-context temporal biases in attention heads and transformer outputs. Using cognitive science methodologies, we analyze attention scores and outputs of the GPT-2 mo…