Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Haowen Wang, Yaxin Du, Jian Yang +9
Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection p…
cs.AI2026
Rubric-Guided Process Reward for Stepwise Model Routing
Shenghao Ye, Yu Guo, Zhengheng Li +2
Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods formulate routing as a sequenti…