3 papers
cs.LG2026
Dual Consensus: Escaping from Spurious Majority in Unsupervised RLVR via Two-Stage Vote Mechanism
Kaixuan Du, Meng Cao, Hang Zhang +3
Current label-free RLVR approaches for large language models (LLMs), such as TTRL and Self-reward, have demonstrated effectiveness in improving the performance of LLMs on complex r…
cs.LG2026
Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables
Andres Potapczynski, Ravi Kiran Selvam, Tatiana Konstantinova +7
In many time series forecasting settings, the target time series is accompanied by exogenous covariates, such as promotions and prices in retail demand; temperature in energy load;…
cs.CL2024
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
Meng Cao, Lei Shu, Lei Yu +4
Reinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences. However, a major challenge arises from the sparsity of thes…