8 citations · 22 across the 40 of their papers we have counts for
4 papers · 1 filter
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
Yongcan Yu, Lingxiao He, Jian Liang +5
Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through…
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
Yongcan Yu, Lingxiao He, Shuo Lu +10
Recent advances in vision-language models (VLMs) reasoning have been largely attributed to the rise of reinforcement Learning (RL), which has shifted the community's focus away fro…
Generative Large-Scale Pre-trained Models for Automated Ad Bidding Optimization
Yu Lei, Jiayang Zhao, Yilei Zhao +4
Modern auto-bidding systems are required to balance overall performance with diverse advertiser goals and real-world constraints, reflecting the dynamic and evolving needs of the i…
Deep Page-Level Interest Network in Reinforcement Learning for Ads Allocation
Guogang Liao, Xiaowen Shi, Ze Wang +5
A mixed list of ads and organic items is usually displayed in feed and how to allocate the limited slots to maximize the overall revenue is a key problem. Meanwhile, modeling user…