11 citations · 21 across the 22 of their papers we have counts for
12 papers · 1 filter
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
Run Wang, Victor J. B. Jung, Philip Wiese +3
On-device tuning of deep neural networks enables long-term adaptation at the edge while preserving data privacy. However, the high computational and memory demands of backpropagati…
Improving Chip Design Enablement for Universities in Europe -- A Position Paper
Lukas Krupp, Ian O'Connor, Luca Benini +3
The semiconductor industry is pivotal to Europe's economy, especially within the industrial and automotive sectors. However, Europe faces a significant shortfall in chip design cap…
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
Gamze İslamoğlu, Luca Bertaccini, Arpan Suravi Prasad +3
Fast and energy-efficient low-bitwidth floating-point (FP) arithmetic is essential for Artificial Intelligence (AI) systems. Microscaling (MX) standardized formats have recently em…
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
Run Wang, Gamze Islamoglu, Andrea Belano +4
While Transformers are dominated by Floating-Point (FP) Matrix-Multiplications, their aggressive acceleration through dedicated hardware or many-core programmable systems has shift…
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
Victor J. B. Jung, Alessio Burrello, Francesco Conti +1
The success of DNNs and their high computational requirements pushed for large codesign efforts aiming at DNN acceleration. Since DNNs can be represented as static computational gr…
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
Sergio Mazzola, Yichao Zhang, Marco Bertuletti +2
As computational paradigms evolve, applications such as attention-based models, wireless telecommunications, and computer vision impose increasingly challenging requirements on com…