52 citations · 82 across the 13 of their papers we have counts for
17 papers · 1 filter
Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
Nishant Subramani
Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come hand-in-hand with increases in s…
The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust
Nishant Subramani, Palash Goyal, Yiwen Song +4
As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good proxy for trust: well-calibrated c…
On the Persistent Effects of Lexicality in Large Language Models
Hammad Rizwan, Muhammad Umair Haider, Nishant Subramani +3
Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these representations is often influen…
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Michael Li, Nishant Subramani
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by me…
Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion
Maitrey Mehta, Nishant Subramani, Zhichao Xu +2
All languages are equal; when it comes to tokenization, some are more equal than others. Tokens are the hidden currency that dictate the cost and latency of access to contemporary…
Personal Information Parroting in Language Models
Nishant Subramani, Kshitish Ghate, Mona Diab
Modern language models (LM) are trained on large scrapes of the Web, containing millions of personal information (PI) instances, many of which LMs memorize, increasing privacy risk…