2 papers
cs.LG2026
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
Tobias Habermann, Michael Mecik, Zhenyu Wang +3
Among hardware accelerators for deep-learning inference, data flow implementations offer low latency and high throughput capabilities. In these architectures, each neuron is mapped…
cs.LG2023
SmoothQuant+: Accurate and Efficient 4-bit Post-Training WeightQuantization for LLM
Jiayi Pan, Chengcan Wang, Kaifu Zheng +3
Large language models (LLMs) have shown remarkable capabilities in various tasks. However their huge model size and the consequent demand for computational and memory resources als…