13 papers · 1 filter
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
Run Wang, Victor J. B. Jung, Philip Wiese +3
On-device tuning of deep neural networks enables long-term adaptation at the edge while preserving data privacy. However, the high computational and memory demands of backpropagati…
Improving Chip Design Enablement for Universities in Europe -- A Position Paper
Lukas Krupp, Ian O'Connor, Luca Benini +3
The semiconductor industry is pivotal to Europe's economy, especially within the industrial and automotive sectors. However, Europe faces a significant shortfall in chip design cap…
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
Serena Curzel, Fabrizio Ferrandi, Leandro Fiorin +15
Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performanc…
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
Gamze İslamoÄlu, Luca Bertaccini, Arpan Suravi Prasad +3
Fast and energy-efficient low-bitwidth floating-point (FP) arithmetic is essential for Artificial Intelligence (AI) systems. Microscaling (MX) standardized formats have recently em…
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
Run Wang, Gamze Islamoglu, Andrea Belano +4
While Transformers are dominated by Floating-Point (FP) Matrix-Multiplications, their aggressive acceleration through dedicated hardware or many-core programmable systems has shift…
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
Severin Bochem, Victor J. B. Jung, Arpan Prasad +2
Contextual Artificial Intelligence (AI) based on emerging Transformer models is predicted to drive the next technology revolution in interactive wearable devices such as new-genera…