32 citations · 176 across the 119 of their papers we have counts for
51 papers · 1 filter
Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
Alex Massucco, Leonardo Del Grande, Marcello Carioni +2
In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretica…
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika +6
The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear min…
Christoffel-DPS: Optimal sensor placement in diffusion posterior sampling for arbitrary distributions
James Rowbottom, Nick Huang, Carola-Bibiane Schönlieb +1
State estimation is a critical task in scientific, engineering and control applications. Since the reliability of reconstructions depends on the number and position of sensors, opt…
Bridging Input Feature Spaces Towards Graph Foundation Models
Moshe Eliasof, Krishna Sri Ipsit Mantri, Beatrice Bevilacqua +2
Unlike vision and language domains, graph learning lacks a shared input space, as input features differ across graph datasets not only in semantics, but also in value ranges and di…
Towards Improved Sentence Representations using Token Graphs
Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Zorah Lähner +1
Obtaining a single-vector representation from a Large Language Model's (LLM) token-level outputs is a critical step for nearly all sentence-level tasks. However, standard pooling m…
Approximation Theory for Lipschitz Continuous Transformers
Takashi Furuya, Davide Murari, Carola-Bibiane Schönlieb
Stability and robustness are critical for deploying Transformers in safety-sensitive settings. A principled way to enforce such behavior is to constrain the model's Lipschitz const…