2 papers
cs.AR2025
Bare-Metal RISC-V + NVDLA SoC for Efficient Deep Learning Inference
Vineet Kumar, Ajay Kumar M, Yike Li +2
This paper presents a novel System-on-Chip (SoC) architecture for accelerating complex deep learning models for edge computing applications through a combination of hardware and so…
cs.AR2025
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
Weiyu Zhou, Zheng Wang, Chao Chen +4
While recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation, centered on high,volume data movem…