activity
20182026
most citedOn the Connection Between Adversarial Robustness and Saliency Map Interpretability

32 citations · 176 across the 119 of their papers we have counts for

collaborators
Showing cs.LGShow all

51 papers · 1 filter

cs.LG2026

Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows

Alex Massucco, Leonardo Del Grande, Marcello Carioni +2

In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretica…

cs.LG2026

Muon is Not That Special: Random or Inverted Spectra Work Just as Well

Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika +6

The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear min…

cs.LG2026

Christoffel-DPS: Optimal sensor placement in diffusion posterior sampling for arbitrary distributions

James Rowbottom, Nick Huang, Carola-Bibiane Schönlieb +1

State estimation is a critical task in scientific, engineering and control applications. Since the reliability of reconstructions depends on the number and position of sensors, opt…

cs.LG2026

Bridging Input Feature Spaces Towards Graph Foundation Models

Moshe Eliasof, Krishna Sri Ipsit Mantri, Beatrice Bevilacqua +2

Unlike vision and language domains, graph learning lacks a shared input space, as input features differ across graph datasets not only in semantics, but also in value ranges and di…

cs.LG2026

Towards Improved Sentence Representations using Token Graphs

Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Zorah Lähner +1

Obtaining a single-vector representation from a Large Language Model's (LLM) token-level outputs is a critical step for nearly all sentence-level tasks. However, standard pooling m…

cs.LG2026

Approximation Theory for Lipschitz Continuous Transformers

Takashi Furuya, Davide Murari, Carola-Bibiane Schönlieb

Stability and robustness are critical for deploying Transformers in safety-sensitive settings. A principled way to enforce such behavior is to constrain the model's Lipschitz const…