7 papers
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement
Hong Qian, Yuanhao Liu, Zihan Zhou +7
While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative s…
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
Ruiqi Xue, Lei Yuan, Kainuo Cheng +2
Learning constraint-satisfying policies from offline data without risky online interaction is crucial for safety-critical decision making. Conventional methods typically learn cost…
Temporal Difference Learning with Constrained Initial Representations
Jiafei Lyu, Jingwen Yang, Zhongjian Qiao +5
Recently, there have been numerous attempts to enhance the sample efficiency of off-policy reinforcement learning (RL) agents when interacting with the environment, including archi…
Cross-Domain Offline Policy Adaptation via Selective Transition Correction
Mengbei Yan, Jiafei Lyu, Shengjie Sun +5
It remains a critical challenge to adapt policies across domains with mismatched dynamics in reinforcement learning (RL). In this paper, we study cross-domain offline RL, where an…
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
Jiafei Lyu, Kang Xu, Jiacheng Xu +6
We consider off-dynamics reinforcement learning (RL) where one needs to transfer policies across different domains with dynamics mismatch. Despite the focus on developing dynamics-…
A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning
Shengjie Sun, Runze Liu, Jiafei Lyu +3
Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often…