22 citations · 113 across the 46 of their papers we have counts for
4 papers · 1 filter
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on…
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
Kangwei Xu, Grace Li Zhang, Ulf Schlichtmann +1
Large Language Models (LLMs) have demonstrated remarkable potential in hardware front-end design using hardware description languages (HDLs). However, their inherent tendency towar…
LiveMind: Low-latency Large Language Models with Simultaneous Inference
Chuangtao Chen, Grace Li Zhang, Xunzhao Yin +3
In this paper, we introduce LiveMind, a novel low-latency inference framework for large language model (LLM) inference which enables LLMs to perform inferences with incomplete user…
Class-Aware Pruning for Efficient Neural Networks
Mengnan Jiang, Jingcun Wang, Amro Eldebiky +4
Deep neural networks (DNNs) have demonstrated remarkable success in various fields. However, the large number of floating-point operations (FLOPs) in DNNs poses challenges for thei…