From the 1 of 11 linked papers with an AI index.
31 citations · 31 across the 1 of their papers we have counts for
11 papers
Improving Autonomous Nano-drones Performance via Automated End-to-End Optimization and Deployment of DNNs
Vlad Niculescu, Lorenzo Lamberti, Francesco Conti +2
The paper presents an automated workflow to train, optimize, and deploy a vision-based CNN (PULP‑Dronet) on an ultra‑low‑power multicore SoC for autonomous navigation of sub‑10 cm…
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
Run Wang, Victor J. B. Jung, Philip Wiese +3
On-device tuning of deep neural networks enables long-term adaptation at the edge while preserving data privacy. However, the high computational and memory demands of backpropagati…
Safe-NEureka: a Hybrid Modular Redundant DNN Accelerator for On-board Satellite AI Processing
Riccardo Tedeschi, Luigi Ghionda, Alessandro Nadalini +5
Low Earth Orbit (LEO) constellations are revolutionizing the space sector, with on-board Artificial Intelligence (AI) becoming pivotal for next-generation satellites. AI accelerati…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +6
Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of…
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
Gamze İslamoÄlu, Luca Bertaccini, Arpan Suravi Prasad +3
Fast and energy-efficient low-bitwidth floating-point (FP) arithmetic is essential for Artificial Intelligence (AI) systems. Microscaling (MX) standardized formats have recently em…
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
Run Wang, Gamze Islamoglu, Andrea Belano +4
While Transformers are dominated by Floating-Point (FP) Matrix-Multiplications, their aggressive acceleration through dedicated hardware or many-core programmable systems has shift…