most citedSTAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar

7 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AR2024

Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity

Cenlin Duan, Jianlei Yang, Yiou Wang +7

Bit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computatio…

cs.AR20247 cited

STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar

Yifeng Zhai, Bing Li, Bonan Yan +1

RRAM crossbars have been studied to construct in-memory accelerators for neural network applications due to their in-situ computing capability. However, prior RRAM-based accelerato…

cs.AR2024

AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology

Rongqing Cong, Wenyang He, Mingxuan Li +5

Large language models (LLMs) with Transformer architectures have become phenomenal in natural language processing, multimodal generative artificial intelligence, and agent-oriented…

cs.AR2023

DDC-PIM: Efficient Algorithm/Architecture Co-design for Doubling Data Capacity of SRAM-based Processing-In-Memory

Cenlin Duan, Jianlei Yang, Xiaolin He +9

Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been d…

cs.AR2023

Fast and reconfigurable sort-in-memory system enabled by memristors

Lianfeng Yu, Yaoyu Tao, Teng Zhang +10

Sorting is fundamental and ubiquitous in modern computing systems. Hardware sorting systems are built based on comparison operations with Von Neumann architecture, but their perfor…