Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
Chuanyang Zheng, Jiankai Sun, Yihang Gao +13
Mixture-of-Experts (MoE) has become a cornerstone in recent state-of-the-art large language models (LLMs). Traditionally, MoE relies on as the router score funct…
cs.CL2025
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
Roland Riachi, Kashif Rasul, Arjun Ashok +5
Recent works have demonstrated the effectiveness of adapting pre-trained language models (LMs) for forecasting time series in the low-data regime. We build upon these findings by a…