3 papers
cs.LG2026
From the Inside Out: Progressive Distribution Refinement for Confidence Calibration
Xizhong Yang, Yinan Xia, Huiming Wang +1
Leveraging the model's internal information as the self-reward signal in Reinforcement Learning (RL) has received extensive attention due to its label-free nature. While prior work…
cs.LG2026
Believe Your Model: Distribution-Guided Confidence Calibration
Xizhong Yang, Haotian Zhang, Huiming Wang +1
Large Reasoning Models have demonstrated remarkable performance with the advancement of test-time scaling techniques, which enhances prediction accuracy by generating multiple cand…
cs.CL2026
Semantic Bridging Domains: Pseudo-Source as Test-Time Connector
Xizhong Yang, Huiming Wang, Ning Xu +1
Distribution shifts between training and testing data are a critical bottleneck limiting the practical utility of models, especially in real-world test-time scenarios. To adapt mod…