1 paper
Xingyu Zhao, Darsh Sharma, Rheeya Uppaal +1
Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy bet…