17 papers
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
Bowen Wang, Chi Zhang, Diyou Shen +3
The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…
Scalable Attention for 5G NR Channel Estimation
Mahdi Abdollahpour, Marco Bertuletti, Yichao Zhang +2
Attention-based neural estimators achieve strong channel-estimation accuracy, but the computational cost of global attention over the time-frequency resource grid grows quadratical…
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
Yinrong Li, Zexin Fu, Yichao Zhang +5
Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchro…
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
Marco Bertuletti, Yichao Zhang, Diyou Shen +3
The upcoming integration of AI in the physical layer (PHY) of 6G radio access networks (RAN) will enable a higher quality of service in challenging transmission scenarios. However,…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +1
Attention accounts for an increasingly dominant fraction of total computation during inference for mixture-of-experts (MoE) models, making efficient acceleration critical. Emerging…
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
Yichao Zhang, Marco Bertuletti, Chi Zhang +5
Shared L1-memory clusters of streamlined instruction processors (processing elements - PEs) are commonly used as building blocks in modern, massively parallel computing architectur…