most citedPushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV20251 cited

Enhancing Audio-Visual Spiking Neural Networks through Semantic-Alignment and Cross-Modal Residual Learning

Xiang He, Dongcheng Zhao, Yiting Dong +3

Humans interpret and perceive the world by integrating sensory information from multiple modalities, such as vision and hearing. Spiking Neural Networks (SNNs), as brain-inspired c…

cs.AR20251 cited

Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA

Jindong Li, Tenglong Li, Guobin Shen +3

The extremely high computational and storage demands of large language models have excluded most edge devices, which were widely used for efficient machine learning, from being via…

cs.NE2025

: Enhanced Information Flow in Spiking Neural Networks with High Hardware Compatibility

Guobin Shen, Jindong Li, Tenglong Li +2

Spiking Neural Networks (SNNs) hold promise for energy-efficient, biologically inspired computing. We identify substantial informatio loss during spike transmission, linked to temp…

cs.CR2024

Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models

Yiting Dong, Guobin Shen, Dongcheng Zhao +2

Large Language Models (LLMs) remain vulnerable to jailbreak attacks that bypass their safety mechanisms. Existing attack methods are fixed or specifically tailored for certain mode…

cs.AR2024

Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines

Jindong Li, Tenglong Li, Guobin Shen +3

Systolic architectures are widely embraced by neural network accelerators for their superior performance in highly parallelized computation. The DSP48E2s serve as dedicated arithme…