6 papers
MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment
Zixuan Yang, Yibo Zhao, Weicong Liu +1
Matching submissions with suitable reviewers at scale is a growing challenge for major venues, yet existing approaches either rely on coarse proxy signals that conflate general rel…
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
Zixuan Yang, Yiqun Chen, Wei Yang +7
Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based metho…
Simply Stabilizing the Loop via Fully Looped Transformer
Rao Fu, Zixuan Yang, Jiankun Zhang +4
Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading a…
Scaling Laws for Online Advertisement Retrieval
Yunli Wang, Zhen Zhang, Zixuan Yang +9
The scaling law is a notable property of neural network models and has significantly propelled the development of large language models. Scaling laws hold great promise in guiding…
Learning Cascade Ranking as One Network
Yunli Wang, Zhen Zhang, Zhiqiang Wang +6
Cascade Ranking is a prevalent architecture in large-scale top-k selection systems like recommendation and advertising platforms. Traditional training methods focus on single-stage…
Adaptive: Adaptive Domain Mining for Fine-grained Domain Adaptation Modeling
Wenxuan Sun, Zixuan Yang, Yunli Wang +8
Advertising systems often face the multi-domain challenge, where data distributions vary significantly across scenarios. Existing domain adaptation methods primarily focus on build…