7 papers
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
Jihwan Oh, Soowon Oh, Murad Aghazada +3
Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identi…
From Belief Entrenchment to Robust Reasoning in LLM Agents
Jihwan Oh, Minchan Jeong, Jongwoo Ko +1
Multi-Agent Debate (MAD) has emerged as a promising inference scaling method for Large Language Model (LLM) reasoning. However, it frequently suffers from belief entrenchment, wher…
Efficient Parametric SVD of Koopman Operator for Stochastic Dynamical Systems
Minchan Jeong, J. Jon Ryu, Se-Young Yun +1
The Koopman operator provides a principled framework for analyzing nonlinear dynamical systems through linear operator theory. Recent advances in dynamic mode decomposition (DMD) h…
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
Jung-hun Kim, Se-Young Yun, Minchan Jeong +3
We study contextual linear bandit problems under feature uncertainty, where the features are noisy and have missing entries. To address the challenges posed by this noise, we analy…
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
Gihun Lee, Minchan Jeong, Yujin Kim +4
While learning to align Large Language Models (LLMs) with human preferences has shown remarkable success, aligning these models to meet the diverse user preferences presents furthe…
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
Yunseon Choi, Sangmin Bae, Seonghyun Ban +6
With the advent of foundation models, prompt tuning has positioned itself as an important technique for directing model behaviors and eliciting desired responses. Prompt tuning reg…