11 papers
Physically-Aware Preemptive Virtual Channels for Deadlock-Free AXI Networks-on-Chip
Lorenzo Leone, Luca Colagrande, Luca Benini
As many-core Systems-on-Chip (SoCs) continue to scale, Networks-on-Chip (NoCs) must sustain increasingly high memory bandwidth while preserving deadlock freedom. In AXI4 systems, p…
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
Luca Colagrande, Lorenzo Leone, Chen Wu +3
The exponential increase in Machine Learning (ML) model size and complexity has driven unprecedented demand for high-performance acceleration systems. As technology scaling enables…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +1
Attention accounts for an increasingly dominant fraction of total computation during inference for mixture-of-experts (MoE) models, making efficient acceleration critical. Emerging…
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
Luca Colagrande, Luca Benini
Large-scale ML accelerators rely on large numbers of PEs, imposing strict bounds on the area and energy budget of each PE. Prior work demonstrates that limited dual-issue capabilit…
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
Luca Colagrande, Lorenzo Leone, Maximilian Coco +2
The growing computational demands of machine learning (ML) workloads have driven the design of ML accelerators aiming at an optimal tradeoff between efficiency and flexibility. A w…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +6
Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of…