3 papers
cs.LG2026
HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
Jia Wei, Zhonghao Zhang, Ping Chen +5
Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning of large language models, yet most variants target dense architectures. Mixture-of-Experts (MoE) models scale p…
cs.LG2025
AWEMixer: Adaptive Wavelet-Enhanced Mixer Network for Long-Term Time Series Forecasting
Qianyang Li, Xingjun Zhang, Peng Tao +3
Forecasting long-term time series in IoT environments remains a significant challenge due to the non-stationary and multi-scale characteristics of sensor signals. Furthermore, erro…
cs.CL2025
Adaptive Rectification Sampling for Test-Time Compute Scaling
Zhendong Tan, Xingjun Zhang, Chaoyi Hu +2
The newly released OpenAI-o1 and DeepSeek-R1 have demonstrated that test-time scaling can significantly improve model performance, especially in complex tasks such as logical reaso…