1 paper
Ruotian Ma, Peisong Wang, Cheng Liu +6
Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs' deep thinking abilities generally require large-scale…