2 papers
cs.AR2026
A Systolic Array Architecture for Nonlinear Activation Functions and Softmax Computation using Chebyshev Polynomials
Benedikt Schaible, Anirudh Suresh Bharadwaj, Ulf Schlichtmann +1
Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware un…
cs.AR2026
Fast Cross-Operator Optimization of Attention Dataflow
Haodong Chang, Hailiang Hu, Zhenrui Wang +5
Attention is a fundamental computational kernel that accounts for the majority of the workload in transformer and LLM computing. Optimizing dataflow is crucial for enhancing both p…