Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Believe Your Model: Distribution-Guided Confidence Calibration
Xizhong Yang, Haotian Zhang, Huiming Wang +1
Large Reasoning Models have demonstrated remarkable performance with the advancement of test-time scaling techniques, which enhances prediction accuracy by generating multiple cand…
cs.LG2025
SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
Jinghui Wang, Shaojie Wang, Yinghan Cui +24
We introduce SeamlessFlow, a server based reinforcement learning (RL) framework that addresses two core challenges in industrial scale RL: (1) decoupling RL training from the compl…