2 papers
cs.AR2026
Shift-Accumulate Attention: Multiplier-Free Query--Key Products for Transformer Decoding
Khubaib Ahmed, Amna Noor, Ahsan Ul haq
Power-of-two (PoT) quantisation turns a multiplication into a bit shift, so far only for the post-softmax attention--value product. The earlier and larger product, S=QK^T, has not…
cs.LG2025
Optimizing Sensory Neurons: Nonlinear Attention Mechanisms for Accelerated Convergence in Permutation-Invariant Neural Networks for Reinforcement Learning
Junaid Muzaffar, Khubaib Ahmed, Ingo Frommholz +2
Training reinforcement learning (RL) agents often requires significant computational resources and prolonged training durations. To address this challenge, we build upon prior work…