5 papers
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
Aleksander Lorenc, Frédéric Berdoz, Joël Mathys +1
Improving the inference efficiency of autoregressive transformers typically means reducing FLOPs per token, usually through approximations that degrade model quality. We introduce…
From Message-Passing to Linearized Graph Sequence Models
Joël Mathys, Basil Rohner, Saku Peltonen +1
Message-passing based approaches form the default backbone of most learning architectures on graph-structured data. However, the rapid progress of modern deep learning architecture…
Learn to Jump: Adaptive Random Walks for Long-Range Propagation through Graph Hierarchies
Joël Mathys, Federico Errica
Message-passing architectures struggle to sufficiently model long-range dependencies in node and graph prediction tasks. We propose a novel approach exploiting hierarchical graph s…
Synthetic Data for Blood Vessel Network Extraction
Joël Mathys, Andreas Plesner, Jorel Elmiger +1
Blood vessel networks in the brain play a crucial role in stroke research, where understanding their topology is essential for analyzing blood flow dynamics. However, extracting de…
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
Niccolò Grillo, Andrea Toccaceli, Joël Mathys +3
Despite incredible progress, many neural architectures fail to properly generalize beyond their training distribution. As such, learning to reason in a correct and generalizable wa…