activity
20242026
collaborators

10 papers

cs.LG2026

Beyond Single Slot: Joint Optimization for Multi-Slot Guaranteed Display Advertising

Zhaoqi Zhang, Jiaming Deng, Miao Xie +5

Guaranteed display advertising is crucial for platform monetization, yet existing methods often operate under a single-slot assumption, limiting their ability to optimize allocatio…

cs.LG2026

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

Yongcan Yu, Lingxiao He, Jian Liang +5

Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through…

cs.AI2026

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation

Shuo Lu, Jianjie Cheng, Yinuo Xu +16

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the…

cs.CL2026

How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1

Yinuo Xu, Shuo Lu, Jianjie Cheng +5

Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. While reinforcement learning (RL) has been shown to improve pe…

cs.LG2025

Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning

Yongcan Yu, Lingxiao He, Shuo Lu +10

Recent advances in vision-language models (VLMs) reasoning have been largely attributed to the rise of reinforcement Learning (RL), which has shifted the community's focus away fro…

cs.LG2025

Generative Large-Scale Pre-trained Models for Automated Ad Bidding Optimization

Yu Lei, Jiayang Zhao, Yilei Zhao +4

Modern auto-bidding systems are required to balance overall performance with diverse advertiser goals and real-world constraints, reflecting the dynamic and evolving needs of the i…