2 papers
cs.CV2024
FLIQS: One-Shot Mixed-Precision Floating-Point and Integer Quantization Search
Jordan Dotzel, Gang Wu, Andrew Li +9
Quantization has become a mainstream compression technique for reducing model size, computational requirements, and energy consumption for modern deep neural networks (DNNs). With…
cs.LG2024
Large Language Models as Optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu +4
Optimization is ubiquitous. While derivative-based algorithms have been powerful tools for various problems, the absence of gradient imposes challenges on many real-world applicati…