3 papers
cs.NI2019
Network-Accelerated Non-Contiguous Memory Transfers
Salvatore Di Girolamo, Konstantin Taranov, Andreas Kurth +7
Applications often communicate data that is non-contiguous in the send- or the receive-buffer, e.g., when exchanging a column of a matrix stored in row-major order. While non-conti…
cs.DC2018
NTX: An Energy-efficient Streaming Accelerator for Floating-point Generalized Reduction Workloads in 22nm FD-SOI
Fabian Schuiki, Michael Schaffner, Luca Benini
Specialized coprocessors for Multiply-Accumulate (MAC) intensive workloads such as Deep Learning are becoming widespread in SoC platforms, from GPUs to mobile SoCs. In this paper w…
eess.IV2017
Hydra: An Accelerator for Real-Time Edge-Aware Permeability Filtering in 65nm CMOS
Manuel Eggimann, Christelle Gloor, Florian Scheidegger +4
Many modern video processing pipelines rely on edge-aware (EA) filtering methods. However, recent high-quality methods are challenging to run in real-time on embedded hardware due…