4 papers
A3 : an Analytical Low-Rank Approximation Framework for Attention
Jeffrey T. H. Wong, Cheng Zhang, Xinye Cao +5
Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-rank approximation offers a promising…
Scaling Laws For Mixed Quantization
Zeyu Cao, Boyang Gu, Cheng Zhang +5
Post-training quantization of Large Language Models (LLMs) has proven effective in reducing the memory and computational requirements for inference. In this study, we focus on a st…
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
Pedro Gimenes, Zeyu Cao, Jeffrey Wong +1
Recent research has shown that LLM performance on reasoning tasks can be enhanced by scaling test-time compute. One promising approach, particularly with decomposable problems, inv…
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
Pedro Gimenes, Yiren Zhao, George Constantinides
Graph Neural Networks (GNNs) have recently gained attention due to their performance on non-Euclidean data. The use of custom hardware architectures proves particularly beneficial…