9 papers · 1 filter
Spectral Conditioning of Attention Improves Transformer Performance
Hemanth Saratchandran, Simon Lucey
We present a theoretical analysis of the Jacobian of an attention block within a transformer, showing that it is governed by the query, key, and value projections that define the a…
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
Jianqiao Zheng, Hemanth Saratchandran, Simon Lucey
Implicit Neural Representations (INRs) have revolutionized continuous signal modeling, yet they struggle to recover fine-grained details within finite training budgets. While empir…
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
Jianqiao Zheng, Cameron Gordon, Yiping Ji +2
Task-agnostic tabular foundation models such as TabPFN have achieved impressive performance on tabular learning tasks, yet the origins of their inductive biases remain poorly under…
Cutting the Skip: Training Residual-Free Transformers
Yiping Ji, James Martens, Jianqiao Zheng +5
Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connectio…
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
Arpit Garg, Hemanth Saratchandran, Ravi Garg +1
Machine unlearning in foundation models (e.g., language and vision transformers) is essential for privacy and safety; however, existing approaches are unstable and unreliable. A wi…
Leaner Transformers: More Heads, Less Depth
Hemanth Saratchandran, Damien Teney, Simon Lucey
Transformers have reshaped machine learning by utilizing attention mechanisms to capture complex patterns in large datasets, leading to significant improvements in performance. Thi…