2 citations · 2 across the 19 of their papers we have counts for
21 papers · 1 filter
Conditioned Initialization for Attention
Hemanth Saratchandran, Simon Lucey
Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their success lies the attention laye…
Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
Damien Teney, Liangze Jiang, Hemanth Saratchandran +1
Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be…
The Quantization Benefits of Residual-Free Transformers
Yiping Ji, Mahalakshmi Sabanayagam, Peyman Moghadam +2
Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantizati…
Preconditioned Attention: Enhancing Efficiency in Transformers
Hemanth Saratchandran
Central to the success of Transformers is the attention block, which effectively models global dependencies among input tokens associated to a dataset. However, we theoretically de…
Spectral Conditioning of Attention Improves Transformer Performance
Hemanth Saratchandran, Simon Lucey
We present a theoretical analysis of the Jacobian of an attention block within a transformer, showing that it is governed by the query, key, and value projections that define the a…
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
Jianqiao Zheng, Hemanth Saratchandran, Simon Lucey
Implicit Neural Representations (INRs) have revolutionized continuous signal modeling, yet they struggle to recover fine-grained details within finite training budgets. While empir…