7 papers
Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning
Feihu Jin, Shipeng Cen, Ying Tan
Fine-tuning large language models (LLMs) achieves strong performance but is often limited by the memory overhead of backpropagation. Zeroth-order (ZO) optimization avoids this over…
Beyond Algorithm Evolution: An LLM-Driven Framework for the Co-Evolution of Swarm Intelligence Optimization Algorithms and Prompts
Shipeng Cen, Ying Tan
The field of automated algorithm design has been advanced by frameworks such as EoH, FunSearch, and Reevo. Yet, their focus on algorithm evolution alone, neglecting the prompts tha…
S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
Wenshuo Wang, Yaomin Shen, Yingjie Tan +1
Spatiotemporal forecasting often relies on computationally intensive models to capture complex dynamics. Knowledge distillation (KD) has emerged as a key technique for creating lig…
Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks
Shipeng Cen, Ying Tan
As optimization problems grow increasingly complex and diverse, advancements in optimization techniques and paradigm innovations hold significant importance. The challenges posed b…
Schrödinger bridge for generative AI: Soft-constrained formulation and convergence analysis
Jin Ma, Ying Tan, Renyuan Xu
Generative AI can be framed as the problem of learning a model that maps simple reference measures into complex data distributions, and it has recently found a strong connection to…
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
Haiquan Qiu, You Wu, Yingjie Tan +2
Loss explosions in training deep neural networks can nullify multi-million dollar training runs. Conventional monitoring metrics like weight and gradient norms are often lagging an…