4 papers
Bridging Quantized Artificial Neural Networks and Neuromorphic Hardware
Zhenhui Chen, Haoran Xu, Yangfan Hu +5
Neuromorphic hardware aims to leverage distributed computing and event-driven circuit design to achieve an energy-efficient AI system. The name "neuromorphic" is derived from its s…
GS-Cache: A GS-Cache Inference Framework for Large-scale Gaussian Splatting Models
Miao Tao, Yuanzhen Zhou, Haoran Xu +10
Rendering large-scale 3D Gaussian Splatting (3DGS) model faces significant challenges in achieving real-time, high-fidelity performance on consumer-grade devices. Fully realizing t…
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Jiecheng Zhou, Ding Tang, Rong Fu +8
The burgeoning computational demands for training large language models (LLMs) necessitate efficient methods, including quantized training, which leverages low-bit arithmetic opera…
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
Haoran Xu, Ziqian Liu, Rong Fu +5
With the evolution of large language models, traditional Transformer models become computationally demanding for lengthy sequences due to the quadratic growth in computation with r…