3 papers
cs.AI2026
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara +4
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution…
cs.LG2026
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
Doo Hwan Hwang, Kee-Eung Kim
Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training. Despite…
cs.CL2024
Zero-Shot Multi-Hop Question Answering via Monte-Carlo Tree Search with Large Language Models
Seongmin Lee, Jaewook Shin, Youngjin Ahn +3
Recent advances in large language models (LLMs) have significantly impacted the domain of multi-hop question answering (MHQA), where systems are required to aggregate information a…