1 paper · 1 filter
Hanmeng Liu, Yiran Ding, Zhizhang Fu +3
Large reasoning models, often post-trained on long chain-of-thought (long CoT) data with reinforcement learning, achieve state-of-the-art performance on mathematical, coding, and d…