4 papers
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
Shijie Zhang, Zheng Xiao, Shiyu Liu +7
Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large language models, but most methods still…
From Item-Only to Query-Item: Query-Conditioned Generative Search with QGS in Quark
Yanglong Song, Zihao Yang, Shuo Meng +6
Generative sequence models have shown strong results in recommendation. Applying them to search ranking is more challenging. Search behavior is inherently query-driven. Each query…
Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning
Shijie Zhang, Xiang Guo, Rujun Guo +4
Building a search relevance model that achieves both low latency and high performance is a long-standing challenge in the search industry. To satisfy the millisecond-level response…
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
Shijie Zhang, Kevin Zhang, Zheyuan Gu +5
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success…