7 papers
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning
Sangmook Lee, Minbeom Kim, Jeonghye Kim +3
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is…
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Dohyung Kim, Minbeom Kim, Jeonghye Kim +3
Reward-maximizing RL methods have shown to be capable of enhancing the reasoning performance of LLMs, but often lead to reduced generation diversity. Recent works address this issu…
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Jeonghye Kim, Xufang Luo, Minbeom Kim +5
Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we…
Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
Sangmook Lee, Dohyung Kim, Hyukhun Koh +2
Recent advances in Large Language Models (LLMs) - particularly model scaling and test-time techniques - have greatly enhanced the reasoning capabilities of language models at the e…
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
Jeonghye Kim, Sojeong Rhee, Minbeom Kim +4
Recent advances in LLM agents have largely built on reasoning backbones like ReAct, which interleave thought and action in complex environments. However, ReAct often produces ungro…
Conditional [MASK] Discrete Diffusion Language Model
Hyukhun Koh, Minha Jhang, Dohyung Kim +2
Although auto-regressive models excel in natural language processing, they often struggle to generate diverse text and provide limited controllability. Non-auto-regressive methods…