Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
Jeesu Jung, Sangkeun Jung
Curriculum learning for training LLMs requires a difficulty signal that aligns with reasoning while remaining scalable and interpretable. We propose a simple premise: tasks that de…
cs.LG2024
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
Hyuk Namgoong, Jeesu Jung, Sangkeun Jung +1
Recent advancements in large language models have heavily relied on the large reward model from reinforcement learning from human feedback for fine-tuning. However, the use of a si…