works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.AR2026

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Bowen Wang, Chi Zhang, Diyou Shen +3

The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…

cs.AR2026

TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks

Marco Bertuletti, Yichao Zhang, Diyou Shen +3

The upcoming integration of AI in the physical layer (PHY) of 6G radio access networks (RAN) will enable a higher quality of service in challenging transmission scenarios. However,…

cs.DC2026

TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link

Yichao Zhang, Marco Bertuletti, Chi Zhang +5

Shared L1-memory clusters of streamlined instruction processors (processing elements - PEs) are commonly used as building blocks in modern, massively parallel computing architectur…

cs.AR2025

TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads

Navaneeth Kunhi Purayil, Diyou Shen, Matteo Perotti +1

The fast evolution of Machine Learning (ML) models requires flexible and efficient hardware solutions as hardwired accelerators face rapid obsolescence. Vector processors are fully…

cs.AR2025

MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster

Sergio Mazzola, Yichao Zhang, Marco Bertuletti +2

As computational paradigms evolve, applications such as attention-based models, wireless telecommunications, and computer vision impose increasingly challenging requirements on com…

cs.AR2025

TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs

Diyou Shen, Yichao Zhang, Marco Bertuletti +1

As computing demand and memory footprint of deep learning applications accelerate, clusters of cores sharing local (L1) multi-banked memory are widely used as key building blocks i…