2 papers
cs.CL2026
ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Models
Shivanshu Kumar, Gopalakrishnan Srinivasan
While the transformer architecture has achieved state-of-the-art performance on natural language processing tasks, these models impose substantial memory and computational overhead…
cs.NE2025
Scaling Equilibrium Propagation to Deeper Neural Network Architectures
Sankar Vinayak Elayedam, Gopalakrishnan Srinivasan
Equilibrium propagation has been proposed as a biologically plausible alternative to the backpropagation algorithm. The local nature of gradient computations, combined with the use…