9 papers
The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning
Joan Bruna
We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}. Each…
Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers
Maya Bechler-Speicher, Gilad Yehudai, Gil Harari +3
Transformers have become a central architecture for graph learning, but their application to graphs requires first choosing a tokenization: a graph-to-token map that determines whi…
Uniform-in-Time Weak Propagation-of-Chaos in Shallow Neural Networks
Margalit Glasgow, Joan Bruna
We consider one-hidden layer neural networks trained in the feature-learning regime using gradient descent, and relate the output of the finite-width network to it…
Geometric Factual Recall in Transformers
Shauli Ravfogel, Gilad Yehudai, Joan Bruna +1
How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of embeddings, requiring parameter…
Geometry and Optimization of Shallow Polynomial Networks
Yossi Arjevani, Joan Bruna, Joe Kileel +2
We study shallow neural networks with monomial activations and output dimension one. The function space for these models can be identified with a set of symmetric tensors with boun…
Axial Neural Networks for Dimension-Free Foundation Models
Hyunsu Kim, Jonggeon Park, Joan Bruna +2
The advent of foundation models in AI has significantly advanced general-purpose learning, enabling remarkable capabilities in zero-shot inference and in-context learning. However,…