6 citations · 6 across the 2 of their papers we have counts for
3 papers
cs.LG2025
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
Yuyang Xu, Yi Cheng, Haochao Ying +5
Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…
cs.CL2025★ 6 cited
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Renjun Hu, Yi Cheng, Libin Meng +4
The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…
cs.IR2025
Behavior Modeling Space Reconstruction for E-Commerce Search
Yejing Wang, Chi Zhang, Xiangyu Zhao +8
Delivering superior search services is crucial for enhancing customer experience and driving revenue growth. Conventionally, search systems model user behaviors by combining user p…