9 papers · 1 filter
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning
Sangmook Lee, Minbeom Kim, Jeonghye Kim +3
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is…
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Jeonghye Kim, Xufang Luo, Minbeom Kim +5
Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we…
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Dohyung Kim, Minbeom Kim, Jeonghye Kim +3
Reward-maximizing RL methods have shown to be capable of enhancing the reasoning performance of LLMs, but often lead to reduced generation diversity. Recent works address this issu…
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
Jeonghye Kim, Sojeong Rhee, Minbeom Kim +4
Recent advances in LLM agents have largely built on reasoning backbones like ReAct, which interleave thought and action in complex environments. However, ReAct often produces ungro…
Drift: Decoding-time Personalized Alignments with Implicit User Preferences
Minbeom Kim, Kang-il Lee, Seongho Joo +3
Personalized alignments for individual users have been a long-standing goal in large language models (LLMs). We introduce Drift, a novel framework that personalizes LLMs at decodin…
Guaranteed Generation from Large Language Models
Minbeom Kim, Thibaut Thonet, Jos Rozen +3
As large language models (LLMs) are increasingly used across various applications, there is a growing need to control text generation to satisfy specific constraints or requirement…