3 papers
cs.SE2026
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
Jianfeng Cai, Jinhua Zhu, Ruopei Sun +5
The rise of reasoning models necessitates large-scale verifiable data, for which programming tasks serve as an ideal source. However, while competitive programming platforms provid…
cs.AI2025
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
Ruopei Sun, Jianfeng Cai, Jinhua Zhu +5
RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction fol…
cs.LG2025
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Jianfeng Cai, Jinhua Zhu, Ruopei Sun +4
Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward m…