2 papers
cs.LG2026
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
Byeonghu Na, Hyungho Na, Yeongmin Kim +4
Large language models (LLMs) are commonly aligned with human preferences using reinforcement learning from human feedback (RLHF). In this method, LLM policies are generally optimiz…
cs.LG2025
Trajectory-Class-Aware Multi-Agent Reinforcement Learning
Hyungho Na, Kwanghyeon Lee, Sumin Lee +1
In the context of multi-agent reinforcement learning, generalization is a challenge to solve various tasks that may require different joint policies or coordination without relying…