3 papers
cs.LG2026
FOGO: Forgetting-aware Orthogonalization Optimizer
Toan Nguyen, Yang Liu, Trung Le +2
We argue that forgetting is not confined to continual learning but is a general optimization phenomenon: during standard training, dominant mini-batch gradients suppress rare but u…
cs.CL2026
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
Aichen Cai, Anmeng Zhang, Anyu Li +66
We introduce JoyAI-LLM Flash, an efficient Mixture-of-Experts (MoE) language model designed to redefine the trade-off between strong performance and token efficiency in the sub-50B…
cs.SI2026
Modeling Trend Dynamics with Variational Neural ODEs for Information Popularity Prediction
Yuchen Wang, Dongpeng Hou, Weikai Jing +3
Predicting the future popularity of information in online social networks is a crucial yet challenging task, due to the complex spatiotemporal dynamics underlying information diffu…