activity
20182025
most citedAn Algorithm-Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers

90 citations · 249 across the 46 of their papers we have counts for

collaborators
Showing cs.ARShow all

19 papers · 1 filter

cs.AR2025★ 1 cited

SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design

Junyi Wu, Chao Fang, Zhongfeng Wang

The growing scale of large language models (LLMs) has intensified demands on computation and memory, making efficient inference a key challenge. While sparsity can reduce these cos…

cs.AR2025

Enable Lightweight and Precision-Scalable Posit/IEEE-754 Arithmetic in RISC-V Cores for Transprecision Computing

Qiong Li, Chao Fang, Longwei Huang +2

While posit format offers superior dynamic range and accuracy for transprecision computing, its adoption in RISC-V processors is hindered by the lack of a unified solution for ligh…

cs.AR2025

FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization

Aotao Wang, Haikuo Shao, Shaobo Ma +1

State Space Models (SSMs), like recent Mamba2, have achieved remarkable performance and received extensive attention. However, deploying Mamba2 on resource-constrained edge devices…

cs.AR2025

AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design

Yanbiao Liang, Huihong Shi, Haikuo Shao +1

Recently, large language models (LLMs) have achieved huge success in the natural language processing (NLP) field, driving a growing demand to extend their deployment from the cloud…

cs.AR2025

An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer

Zhengke Li, Wendong Mao, Siyu Zhang +2

Recently, large models, such as Vision Transformer and BERT, have garnered significant attention due to their exceptional performance. However, their extensive computational requir…

cs.AR2024★ 15 cited

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

Chao Fang, Man Shi, Robin Geens +3

The widely-used, weight-only quantized large language models (LLMs), which leverage low-bit integer (INT) weights and retain floating-point (FP) activations, reduce storage require…