1 paper · 1 filter
Xumeng Wen, Zihan Liu, Shun Zheng +9
Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant int…