9 papers · 1 filter
A Mechanistic Analysis of Looped Reasoning Language Models
Hugh Blayney, Álvaro Arroyo, Johan Obando-Ceron +4
Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM's layers in the latent di…
Can Graph Foundation Models Generalize Over Architecture?
Benjamin Gutteridge, Michael Bronstein, Xiaowen Dong
Graph foundation models (GFMs) have recently attracted interest due to the promise of graph neural network (GNN) architectures that generalize zero-shot across graphs of arbitrary…
gLSTM: Mitigating Over-Squashing by Increasing Storage Capacity
Hugh Blayney, Álvaro Arroyo, Xiaowen Dong +1
Graph Neural Networks (GNNs) leverage the graph structure to transmit information between nodes, typically through the message-passing mechanism. While these models have found a wi…
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
Enrique Queipo-de-Llano, Álvaro Arroyo, Federico Barbero +4
Attention sinks and compression valleys have attracted significant attention as two puzzling phenomena in large language models, but have been studied in isolation. In this work, w…
Mathematical Foundations of Geometric Deep Learning
Haitz Sáez de Ocáriz Borde, Michael Bronstein
We review the key mathematical concepts necessary for studying Geometric Deep Learning.
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
Huidong Liang, Haitz Sáez de Ocáriz Borde, Baskaran Sripathmanathan +2
Long-range dependencies are critical for effective graph representation learning, yet most existing datasets focus on small graphs tailored to inductive tasks, offering limited ins…