3 papers
cs.CV2026
Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank
Xiangyong Chen, Xiaochuan Lin, Haoran Liu +3
Recent advances in reasoning-induced image quality assessment (IQA) have demonstrated the power of reinforcement learning to rank (RL2R) for training vision-language models (VLMs)…
cs.LG2026
Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning
Shijie Zhang, Xiang Guo, Rujun Guo +4
Building a search relevance model that achieves both low latency and high performance is a long-standing challenge in the search industry. To satisfy the millisecond-level response…
cs.LG2026
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
Shijie Zhang, Kevin Zhang, Zheyuan Gu +5
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success…