Showing 2025Show all
2 papers · 1 filter
cs.AI2025
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
Sunzhu Li, Zhiyu Lin, Shuling Yang +2
Large Reasoning Models (LRMs) are powerful, but they still suffer from inefficient and off-target reasoning. Currently, training-free methods are limited to either rigid heuristics…
cs.LG2025
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Yang Zhou, Sunzhu Li, Shunyu Liu +11
Recent advances in Large Language Models (LLMs) have underscored the potential of Reinforcement Learning (RL) to facilitate the emergence of reasoning capabilities. Despite the enc…