2 papers
cs.CL2026
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning
Dongxu Zhang, Yujun Wu, Yiding Sun +5
While Chain-of-Thought (CoT) prompting empowers Large Language Models (LLMs), ensuring reasoning reliability remains an open challenge. Contrary to the prevailing cascading failure…
cs.LG2025
Uncertainty-aware Reward Design Process
Yang Yang, Xiaolu Zhou, Bosong Ding +1
Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging process due to the inefficiencies and inconsistencies inherent in…