4 papers
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
Ye Qiao, Yian Wang, Zhiheng Chen +2
Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-widths (e.g., 8/4/3-bit), of…
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
Ye Qiao, Zhiheng Chen, Yifan Zhang +2
With the emergence of wearable devices and other embedded systems, deploying large language models (LLMs) on edge platforms has become an urgent need. However, this is challenging…
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
Ye Qiao, Zhiheng Chen, Yian Wang +3
Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and…
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
Ye Qiao, Zhiheng Chen, Yifan Zhang +2
Deploying large language models (LLMs) on edge platforms is challenged by their high computational and memory demands. Although recent low-bit quantization methods (e.g., BitNet, D…