collaborators

5 papers

cs.CL2026

A3 : an Analytical Low-Rank Approximation Framework for Attention

Jeffrey T. H. Wong, Cheng Zhang, Xinye Cao +5

Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-rank approximation offers a promising…

cs.LG2025

An Improved Template for Approximate Computing

Morteza Rezaalipour, Francesco Costa, Marco Biasion +3

Deploying neural networks on edge devices entails a careful balance between the energy required for inference and the accuracy of the resulting classification. One technique for na…

cs.LG2025

BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration

Yuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai +4

Large language models (LLMs) have demonstrated remarkable performance across various machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders thei…

cs.AR2025

ATHEENA: A Toolflow for Hardware Early-Exit Network Automation

Benjamin Biggs, Christos-Savvas Bouganis, George A. Constantinides

The continued need for improvements in accuracy, throughput, and efficiency of Deep Neural Networks has resulted in a multitude of methods that make the most of custom architecture…

cs.LG2025

QERA: an Analytical Framework for Quantization Error Reconstruction

Cheng Zhang, Jeffrey T. H. Wong, Can Xiao +2

The growing number of parameters and computational demands of large language models (LLMs) present significant challenges for their efficient deployment. Recently, there is an incr…