2 papers
cs.AR2026
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation
Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura +4
Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring mon…
cs.LG2024
Towards Accurate and Efficient Sub-8-Bit Integer Training
Wenjin Guo, Donglai Liu, Weiying Xie +7
Neural network training is a memory- and compute-intensive task. Quantization, which enables low-bitwidth formats in training, can significantly mitigate the workload. To reduce qu…