2 papers
cs.LG2026
Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
Yimeng Ye, Shuang Chen, Wenxuan Huang +8
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. Th…
cs.LG2026
PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search
Tianhao Qian, Jiayu Chen, Lixu Wang
LLM-guided search is usually adopted to solve complex tasks by ranking and pruning top- candidates based on evaluator scores. However, irreducible bias still exists even if popu…