2 papers
cs.LG2026
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
Yu Huang, Zihua Zhao, Zhaoxin Huan +9
The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical d…
cs.SE2026
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
Zeming Dong, Qiang Hu, Xiaofei Xie +4
Pre-trained code models lead the era of code intelligence, with multiple models designed with impressive performance. However, one important problem, data augmentation for code dat…