collaborators

8 papers

cs.AR2026

Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding

Patrick Iff, Tommaso Bonato, Maciej Besta +2

Transformer-based large language models are increasingly constrained by data movement as communication bandwidth drops sharply beyond the chip boundary. Wafer-scale integration usi…

cs.AR2026

VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration

Max Wipfli, Gamze İslamoğlu, Navaneeth Kunhi Purayil +2

Compared to the first generation of deep neural networks, dominated by regular, compute-intensive kernels such as matrix multiplications (MatMuls) and convolutions, modern decoder-…

cs.DC2025

Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators

Aofeng Shen, Chi Zhang, Yakup Budanaz +4

Tile-based many-Processing Element (PE) accelerators can achieve competitive performance on General Matrix Multiplication (GEMM), but they are extremely hard to program, as their o…

cs.PF2025

PerfDojo: Automated ML Library Generation for Heterogeneous Architectures

Andrei Ivanov, Siyuan Shen, Gioele Gottardo +5

The increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a signifi…

cs.AR2025

RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures

Patrick Iff, Benigna Bruggmann, Blaise Morel +3

Chiplet architectures are on the rise as they promise to overcome the scaling challenges of monolithic chips. A key component of such architectures is an efficient inter-chiplet in…

cs.AR2025

Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs

Cyril Koenig, Enrico Zelioli, Luca Benini

Embedded heterogeneous systems-on-chip (SoCs) rely on domain-specific hardware accelerators to improve performance and energy efficiency. In particular, programmable multi-core acc…