2 papers
cs.LG2026
Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
Yirong Zeng, Yufei Liu, Xiao Ding +9
A central belief in scaling reinforcement learning with verifiable rewards for instruction following (IF) tasks is that, a diverse mixture of verifiable hard and unverifiable soft…
cs.LG2025
Reactivation: Empirical NTK Dynamics Under Task Shifts
Yuzhi Liu, Zixuan Chen, Zirui Zhang +2
The Neural Tangent Kernel (NTK) offers a powerful tool to study the functional dynamics of neural networks. In the so-called lazy, or kernel regime, the NTK remains static during t…