2 papers
cs.CL2025
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
Bolian Li, Yanran Wu, Xinyu Luo +1
Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where…
cs.LG2025
Stacey: Promoting Stochastic Steepest Descent via Accelerated -Smooth Nonconvex Optimization
Xinyu Luo, Cedar Site Bai, Bolian Li +3
While popular optimization methods such as SGD, AdamW, and Lion depend on steepest descent updates in either or norms, there remains a critical gap in handli…