4 citations · 4 across the 16 of their papers we have counts for
9 papers · 1 filter
Tracking Representation Dynamics in Large Language Models with Persistent Homology
Naman Malhotra, Jay Ambadkar, Abhinav Gupta +4
Large language models are commonly aligned through supervised fine-tuning, yet little is known about how their internal representations evolve during this process. We study alignme…
Topological Signatures of Grokking
Yifan Tang, Qiquan Wang, Inés GarcÃa-Redondo +1
We study the grokking phenomenon through the lens of topology. Using persistent homology on point clouds derived from the embedding matrices of a range of models trained on modular…
Feature Starvation as Geometric Instability in Sparse Autoencoders
Faris Chaudhry, Keisuke Yano, Anthea Monod
Sparse autoencoders (SAEs) are used to disentangle the dense, polysemantic internal representations of large language models (LLMs) into interpretable, monosemantic concepts. Howev…
The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
Aideen Fay, Inés GarcÃa-Redondo, Qiquan Wang +2
Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlooks the high-dimensional, relational, an…
Breaking Symmetry Bottlenecks in GNN Readouts
Mouad Talhi, Arne Wolf, Anthea Monod
Graph neural networks (GNNs) are widely used for learning on structured data, yet their ability to distinguish non-isomorphic graphs is fundamentally limited. These limitations are…
Riemannian Neural Optimal Transport
Alessandro Micheli, Yueqi Cao, Anthea Monod +1
Computational optimal transport (OT) offers a principled framework for generative modeling. Neural OT methods, which use neural networks to learn an OT map (or potential) from data…