2 papers
cs.AI2026
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
Dao Tran, Duc Anh Le, Ngoc Luu +3
Test-time scaling improves language model reasoning by spending additional compute to explore multiple solution trajectories. The key challenge is to maximize accuracy while minimi…
cs.AI2026
Selective Off-Policy Reference Tuning with Plan Guidance
Duc Anh Le, Tien-Phat Nguyen, Thien Huu Nguyen +2
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those fa…