9 papers
ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
Róisín Luo
Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Exis…
Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches, and Certifiable Robustness
RóisÃn Luo, James McDermott, Colm O'Riordan
Lipschitz continuity is a fundamental property of neural networks that characterizes their sensitivity to input perturbations. It plays a pivotal role in deep learning, governing \…
Principles of Lipschitz continuity in neural networks
RóisÃn Luo
Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence. Yet, despite t…
A Stochastic--Geometric Theory of Scaling Laws in Grokking
RóisÃn Luo, Christian Gagné, Jonas Ngnawé +2
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged de…
Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition
RóisÃn Luo, James McDermott, Colm O'Riordan
Perturbation robustness evaluates the vulnerabilities of models, arising from a variety of perturbations, such as data corruptions and adversarial attacks. Understanding the mechan…
Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization
RóisÃn Luo, Alexandru Drimbarean, James McDermott +1
This paper explores a novel paradigm in low-bit (i.e. 4-bits or lower) quantization, differing from existing state-of-the-art methods, by framing optimal quantization as an archite…