Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Thinker: Learning to Think Fast and Slow
Stephen Chung, Wenyu Du, Jie Fu
Recent studies show that the reasoning capabilities of Large Language Models (LLMs) can be improved by applying Reinforcement Learning (RL) to question-answering (QA) tasks in area…
cs.CL2025
Learning from Peers in Reasoning Models
Tongxu Luo, Wenyu Du, Jiaxi Bi +5
Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning process sta…
cs.CL2025
Learning from Failures in Multi-Attempt Reinforcement Learning
Stephen Chung, Wenyu Du, Jie Fu
Recent advancements in reinforcement learning (RL) for large language models (LLMs), exemplified by DeepSeek R1, have shown that even a simple question-answering task can substanti…