From the 1 of 11 linked papers with an AI index.
11 papers
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation…
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
Jonas Rohweder, Subhabrata Dutta, Iryna Gurevych
The paper argues that phenomena such as induction heads, function vectors, and the Hydra effect in Transformer language models can be explained by hierarchical latent structures in…
The FIL Hypothesis: Inductive Biases Help with Kernel Engineering
Nikolai Rozanov, Subhabrata Dutta, Preslav Nakov +1
The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowledge, has become a dominant p…
Patches of Nonlinearity: Instruction Vectors in Large Language Models
Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta +1
Despite the recent success of instruction-tuned language models and their ubiquitous usage, very little is known of how models process instructions internally. In this work, we add…
Expert Preference-based Evaluation of Automated Related Work Generation
Furkan Åahinuç, Subhabrata Dutta, Iryna Gurevych
Expert domain writing, such as scientific writing, typically demands extensive domain knowledge. Although large language models (LLMs) show promising potential in this task, evalua…
Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery
Alireza Bayat Makou, Jingcheng Niu, Subhabrata Dutta +1
Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly interpreted as evidence of d…