12 papers
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
Qitao Tan, Jun Liu, Zheng Zhan +6
Large language models (LLMs) excel across various tasks, but standard first-order (FO) fine-tuning demands considerable memory, significantly limiting real-world deployment. Recent…
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
Yang Xiao, Gen Li, Kaiyuan Deng +5
Training-free acceleration has emerged as an advanced research area in video generation based on diffusion models. The redundancy of latents in diffusion model inference provides a…
Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
Qitao Tan, Sung-En Chang, Rui Xia +10
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising…
Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation
Zhenglun Kong, Zheng Zhan, Shiyue Hou +10
Large language models (LLMs) have shown remarkable promise but remain challenging to continually improve through traditional finetuning, particularly when integrating capabilities…
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
Zheng Zhan, Yushu Wu, Yifan Gong +7
The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. H…
Search for Efficient Large Language Models
Xuan Shen, Pu Zhao, Yifan Gong +7
Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and disti…