gradient decorrelation 1hierarchical latent structures 1mechanistic interpretability 1synthetic data 1transformer models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
Jonas Rohweder, Subhabrata Dutta, Iryna Gurevych
The paper argues that phenomena such as induction heads, function vectors, and the Hydra effect in Transformer language models can be explained by hierarchical latent structures in…
cs.CL2026
Patches of Nonlinearity: Instruction Vectors in Large Language Models
Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta +1
Despite the recent success of instruction-tuned language models and their ubiquitous usage, very little is known of how models process instructions internally. In this work, we add…
cs.AI2025
SwiftSolve: A Self-Iterative, Complexity-Aware Multi-Agent Framework for Competitive Programming
Adhyayan Veer Singh, Aaron Shen, Brian Law +4
Correctness alone is insufficient: LLM-generated programs frequently satisfy unit tests while violating contest time or memory budgets. We present SwiftSolve, a complexity-aware mu…