1 paper · 1 filter
Qipan Wang, Zhe Zhang, Shuangchen Li +5
The success of large language models LLMs amplifies the need for highthroughput energyefficient inference at scale. 3DDRAMbased accelerators provide high memory bandwidth and there…