works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AR2026

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Bowen Wang, Chi Zhang, Diyou Shen +3

The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…

cs.AR2026

VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration

Max Wipfli, Gamze İslamoğlu, Navaneeth Kunhi Purayil +2

Compared to the first generation of deep neural networks, dominated by regular, compute-intensive kernels such as matrix multiplications (MatMuls) and convolutions, modern decoder-…

cs.ET2025

CMOS 2.0 -- Redefining the Future of Scaling

Moritz Brunion, Navaneeth Kunhi Purayil, Francesco Dell'Atti +5

We propose to revisit the functional scaling paradigm by capitalizing on two recent developments in advanced chip manufacturing, namely 3D wafer bonding and backside processing. Th…

cs.AR2025

TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads

Navaneeth Kunhi Purayil, Diyou Shen, Matteo Perotti +1

The fast evolution of Machine Learning (ML) models requires flexible and efficient hardware solutions as hardwired accelerators face rapid obsolescence. Vector processors are fully…

cs.AR2025

AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors

Navaneeth Kunhi Purayil, Matteo Perotti, Tim Fischer +1

The ever-growing scale of data parallelism in today's HPC and ML applications presents a big challenge for computing architectures' energy efficiency and performance. Vector proces…