2 papers
cs.AI2024
LiveMind: Low-latency Large Language Models with Simultaneous Inference
Chuangtao Chen, Grace Li Zhang, Xunzhao Yin +3
In this paper, we introduce LiveMind, a novel low-latency inference framework for large language model (LLM) inference which enables LLMs to perform inferences with incomplete user…
cs.AI2024
Class-Aware Pruning for Efficient Neural Networks
Mengnan Jiang, Jingcun Wang, Amro Eldebiky +4
Deep neural networks (DNNs) have demonstrated remarkable success in various fields. However, the large number of floating-point operations (FLOPs) in DNNs poses challenges for thei…