2 papers
cs.LG2025
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
Cong Xu, Wenbin Liang, Mo Yu +7
The rapid scaling of models has led to prohibitively high training and fine-tuning costs. A major factor accounting for memory consumption is the widespread use of stateful optimiz…
cs.IR2024
Are LLM-based Recommenders Already the Best? Simple Scaled Cross-entropy Unleashes the Potential of Traditional Sequential Recommenders
Cong Xu, Zhangchi Zhu, Mo Yu +3
Large language models (LLMs) have been garnering increasing attention in the recommendation community. Some studies have observed that LLMs, when fine-tuned by the cross-entropy (C…