3 papers
cs.LG2026
Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
Aleksandros Sobczyk, Gioele Gottardo, Christos K. Matzoros +4
Linear attention has emerged as a cornerstone for efficient long-context architectures, as evidenced by its integration into state-of-the-art open-source models including Qwen3.5/3…
cs.DC2026
Parallel Scan on Ascend AI Accelerators
BartÅomiej Wróblewski, Gioele Gottardo, Anastasios Zouzias
We design and implement parallel prefix sum (scan) algorithms using Ascend AI accelerators. Ascend accelerators feature specialized computing units: the cube units for efficient ma…
cs.PF2025
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
Andrei Ivanov, Siyuan Shen, Gioele Gottardo +5
The increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a signifi…