1 citations · 1 across the 3 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2024
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi +27
Monolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications. Training, serving, and maintaining monolithic LLMs at scale, however, r…
cs.AR2021
Capstan: A Vector RDA for Sparsity
Alexander Rucker, Matthew Vilim, Tian Zhao +3
This paper proposes Capstan: a scalable, parallel-patterns-based, reconfigurable dataflow accelerator (RDA) for sparse and dense tensor applications. Instead of designing for one a…