4 papers
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
George Karfakis, Faraz Tahmasebi, Binglu Chen +5
RAPID-LLM is a unified performance modeling framework for distributed large language model (LLM) training and inference on GPU clusters, without relying on deployment-specific trac…
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
Faraz Tahmasebi, Michael Pelluer, Hyoukjun Kwon
The computation and memory costs of large language models kept increasing over last decade, which reached over the scale of 1T parameters. To address the challenges from the large…
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
Faraz Tahmasebi, Yian Wang, Benji Y. H. Huang +1
Recent research has shown that large language models (LLMs) can utilize low-precision floating point (FP) quantization to deliver high efficiency while maintaining original model a…
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
Haocheng Xu, Faraz Tahmasebi, Ye Qiao +3
Recent innovations in Transformer-based large language models have significantly advanced the field of general-purpose neural language understanding and generation. With billions o…