From the 1 of 4 linked papers with an AI index.
2 citations · 2 across the 3 of their papers we have counts for
3 papers · 1 filter
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
Bowen Wang, Chi Zhang, Diyou Shen +3
The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
Yinrong Li, Zexin Fu, Yichao Zhang +5
Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchro…
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
Bowen Wang, Marco Bertuletti, Yichao Zhang +2
Attention-based models demand flexible hardware to manage diverse kernels with varying arithmetic intensities and memory access patterns. Large clusters with shared L1 memory, a co…