4 papers
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
Vasileios Titopoulos, Kosmas Alexandridis, Giorgos Dimitrakopoulos
Attention is a core operation in numerous machine learning and artificial intelligence models. This work focuses on the acceleration of attention kernel using FlashAttention algori…
FLASH-D: FlashAttention with Hidden Softmax Division
Kosmas Alexandridis, Vasileios Titopoulos, Giorgos Dimitrakopoulos
The transformer's attention mechanism has revolutionized AI and machine learning, with its efficient computation being crucial to its performance. However, calculating attention in…
Online Alignment and Addition in Multi-Term Floating-Point Adders
Kosmas Alexandridis, Giorgos Dimitrakopoulos
Multi-term floating-point addition appears in vector dot-product computations, matrix multiplications, and other forms of floating-point data aggregation. A critical step in multi-…
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
Kosmas Alexandridis, Christodoulos Peltekis, Dionysios Filippas +1
The widespread adoption of machine learning algorithms necessitates hardware acceleration to ensure efficient performance. This acceleration relies on custom matrix engines that op…