3 papers
cs.CL2026
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
Jiashu Yao, Heyan Huang, Daiqing Wu +2
To encourage diverse exploration in reinforcement learning (RL) for large language models (LLMs) without compromising accuracy, we propose Policy Split, a novel paradigm that bifur…
cs.CV2025
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
Zhangquan Chen, Ruihui Zhao, Chuwei Luo +4
Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods hav…
cs.CL2025
Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
Jiashu Yao, Heyan Huang, Shuang Zeng +6
Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on compl…