5 papers
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Yujin Kim, Namgyu Ho, Sangmin Hwang +7
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt th…
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
Jihwan Oh, Soowon Oh, Murad Aghazada +3
Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identi…
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
Yongjin Yang, Sinjae Kang, Juyong Lee +3
Training large language model (LLM) agents to acquire necessary skills and perform diverse tasks within an environment is gaining interest as a means to enable open-endedness. Howe…
Self-Training Elicits Concise Reasoning in Large Language Models
Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim +3
Chain-of-thought (CoT) reasoning has enabled large language models (LLMs) to utilize additional computation through intermediate tokens to solve complex tasks. However, we posit th…
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
Gihun Lee, Minchan Jeong, Yujin Kim +4
While learning to align Large Language Models (LLMs) with human preferences has shown remarkable success, aligning these models to meet the diverse user preferences presents furthe…