activity
20242026
most citedOpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.ARShow all

8 papers · 1 filter

cs.AR2026

A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access

Xiaoling Yi, Ryan Antonio, Yunhao Deng +4

Achieving high compute utilization across a wide range of AI workloads is crucial for the efficiency of versatile DNN accelerators. This paper presents the Voltra chip and its util…

cs.AR2025

Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement

Yunhao Deng, Fanchen Kong, Xiaoling Yi +2

The growing disparity between computational power and on-chip communication bandwidth is a critical bottleneck in modern Systems-on-Chip (SoCs), especially for data-parallel worklo…

cs.AR2025

Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration

Stef Cuyckens, Xiaoling Yi, Robin Geens +4

Emerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microsc…

cs.AR2025

An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems

Ryan Albert Antonio, Joren Dumoulin, Xiaoling Yi +4

Heterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movem…

cs.AR2025

XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs

Fanchen Kong, Yunhao Deng, Xiaoling Yi +2

As modern AI workloads increasingly rely on heterogeneous accelerators, ensuring high-bandwidth and layout-flexible data movements between accelerator memories has become a pressin…

cs.AR2025

Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning

Stef Cuyckens, Xiaoling Yi, Nitish Satya Murthy +2

Autonomous robots require efficient on-device learning to adapt to new environments without cloud dependency. For this edge training, Microscaling (MX) data types offer a promising…