1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.DC2024
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
Zongwu Wang, Fangxin Liu, Mingshuai Li +1
Efficient parallelization of Large Language Models (LLMs) with long sequences is essential but challenging due to their significant computational and memory demands, particularly s…
cs.AR2021★ 1 cited
SME: ReRAM-based Sparse-Multiplication-Engine to Squeeze-Out Bit Sparsity of Neural Network
Fangxin Liu, Wenbo Zhao, Yilong Zhao +6
Resistive Random-Access-Memory (ReRAM) crossbar is a promising technique for deep neural network (DNN) accelerators, thanks to its in-memory and in-situ analog computing abilities…