2 papers
cs.AI2026
Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations
Bochao Liu, Zhipeng Qian, Yang Zhao +10
Operating and maintaining (O&M) large-scale online engine systems (eg, search, recommendation and advertising) demands substantial human effort for release monitoring, alert respon…
cs.CL2026
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
Youliang Yuan, Qiuyang Mang, Jingbang Chen +7
In this paper, we observe that current models are susceptible to reward hacking, leading to a substantial overestimation of a model's reasoning ability. This is evidenced by a high…