5 papers
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
Jihwan Oh, Soowon Oh, Murad Aghazada +3
Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identi…
MERIT Feedback Elicits Better Bargaining in LLM Negotiators
Jihwan Oh, Murad Aghazada, Yooju Shin +2
Bargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic…
LLM Agents for Bargaining with Utility-based Feedback
Jihwan Oh
Bargaining, a critical aspect of real-world interactions, presents challenges for large language models (LLMs) due to limitations in strategic depth and adaptation to complex human…
From Belief Entrenchment to Robust Reasoning in LLM Agents
Jihwan Oh, Minchan Jeong, Jongwoo Ko +1
Multi-Agent Debate (MAD) has emerged as a promising inference scaling method for Large Language Model (LLM) reasoning. However, it frequently suffers from belief entrenchment, wher…
Diffusion-based Episodes Augmentation for Offline Multi-Agent Reinforcement Learning
Jihwan Oh, Sungnyun Kim, Gahee Kim +2
Offline multi-agent reinforcement learning (MARL) is increasingly recognized as crucial for effectively deploying RL algorithms in environments where real-time interaction is impra…