3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Towards Interpreting Language Models: A Case Study in Multi-Hop Reasoning
Mansi Sakarvadia
Answering multi-hop reasoning questions requires retrieving and synthesizing information from diverse sources. Language models (LMs) struggle to perform such reasoning consistently…
cs.LG2024★ 1 cited
Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
Nathaniel Hudson, J. Gregory Pauloski, Matt Baughman +13
Deep learning methods are transforming research, enabling new techniques, and ultimately leading to new discoveries. As the demand for more capable AI models continues to grow, we…
cs.CL2023★ 3 cited
Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism
Mansi Sakarvadia, Arham Khan, Aswathy Ajith +5
Transformer-based Large Language Models (LLMs) are the state-of-the-art for natural language tasks. Recent work has attempted to decode, by reverse engineering the role of linear l…