3 papers
cs.AI2025
ChipGPT: How far are we from natural language hardware design
Kaiyan Chang, Ying Wang, Haimeng Ren +5
As large language models (LLMs) like ChatGPT exhibited unprecedented machine intelligence, it also shows great performance in assisting hardware engineers to realize higher-efficie…
cs.AR2025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
Lian Liu, Shixin Zhao, Bing Li +6
The billion-scale Large Language Models (LLMs) need deployment on expensive server-grade GPUs with large-storage HBMs and abundant computation capability. As LLM-assisted services…
cs.AR2024
COMET: Towards Partical W4A4KV4 LLMs Serving
Lian Liu, Haimeng Ren, Long Cheng +6
Quantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalen…