3 papers
cs.CL2026
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
Aryan Sharma, Cutter Dawes, Shivam Raval
When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residu…
cs.CL2026
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
Cutter Dawes, Aryan Sharma, Angelos Ioannis Lagos +1
Representing and navigating hierarchy is a fundamental primitive of reasoning. Large language models have demonstrated proficiency in a wide variety of tasks requiring hierarchical…
cs.LG2025
A Group Theoretic Analysis of the Symmetries Underlying Base Addition and Their Learnability by Neural Networks
Cutter Dawes, Simon Segert, Kamesh Krishnamurthy +1
A major challenge in the use of neural networks both for modeling human cognitive function and for artificial intelligence is the design of systems with the capacity to efficiently…