2 papers
cs.AR2025
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
Mingqiang Huang, Ao Shen, Kai Li +4
The rapid advancements in artificial intelligence (AI), particularly the Large Language Models (LLMs), have profoundly affected our daily work and communication forms. However, it…
cs.AR2025
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
Sixiao Huang, Tintin Wang, Ang Li +5
Large language models (LLMs) are both storage-intensive and computation-intensive, posing significant challenges when deployed on resource-constrained hardware. As linear layers in…