3 papers
cs.LG2026
Orthogonal Self-Attention
Leo Zhang, James Martens
Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, re…
cs.LG2025
Cutting the Skip: Training Residual-Free Transformers
Yiping Ji, James Martens, Jianqiao Zheng +5
Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connectio…
math.AP2025
Discovery of Unstable Singularities
Yongji Wang, Mehdi Bennani, James Martens +19
Whether singularities can form in fluids remains a foundational unanswered question in mathematics. This phenomenon occurs when solutions to governing equations, such as the 3D Eul…