5 papers
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
Keru Chen, Jun Luo, Sen Lin +4
Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…
LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas
Yun Hua, Shang Gao, Wenhao Li +5
In multi-agent reinforcement learning, each agent acts to maximize its individual accumulated rewards. Nevertheless, individual accumulated rewards could not fully reflect how othe…
Learning Virtual Machine Scheduling in Cloud Computing through Language Agents
JieHao Wu, Ziwei Wang, Junjie Sheng +3
In cloud services, virtual machine (VM) scheduling is a typical Online Dynamic Multidimensional Bin Packing (ODMBP) problem, characterized by large-scale complexity and fluctuating…
Online Finetuning Decision Transformers with Pure RL Gradients
Junkai Luo, Yinglun Zhu
Decision Transformers (DTs) have emerged as a powerful framework for sequential decision making by formulating offline reinforcement learning (RL) as a sequence modeling problem. H…
Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents
Yun Hua, Haosheng Chen, Shiqin Wang +3
Large Language Models (LLMs) show strong collaborative performance in multi-agent systems with predefined roles and workflows. However, in open-ended environments lacking coordinat…