works on

From the 1 of 37 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.ARShow all

6 papers · 1 filter

cs.AR2026

FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control

Tenglong Li, Jindong Li, Guobin Shen +3

Spiking Neural Networks (SNNs) offer a biologically plausible learning mechanism through synaptic plasticity, enabling unsupervised adaptation without the computational overhead of…

cs.AR2026

FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture

Tenglong Li, Jindong Li, Guobin Shen +3

Spiking Neural Networks (SNNs), with brain-inspired structure using discrete spikes instead of continuous activations, are gaining attention for their efficient processing on neuro…

cs.AR2025

Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA

Jindong Li, Tenglong Li, Ruiqi Chen +4

Deploying large language models (LLMs) on embedded devices remains a significant research challenge due to the high computational and memory demands of LLMs and the limited hardwar…

cs.AR2025

FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture

Tenglong Li, Jindong Li, Guobin Shen +3

Spiking transformers are emerging as a promising architecture that combines the energy efficiency of Spiking Neural Networks (SNNs) with the powerful attention mechanisms of transf…

cs.AR2025

Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA

Jindong Li, Tenglong Li, Guobin Shen +3

The extremely high computational and storage demands of large language models have excluded most edge devices, which were widely used for efficient machine learning, from being via…

cs.AR2024

Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines

Jindong Li, Tenglong Li, Guobin Shen +3

Systolic architectures are widely embraced by neural network accelerators for their superior performance in highly parallelized computation. The DSP48E2s serve as dedicated arithme…