activity
20182023
most citedSMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

103 citations · 407 across the 41 of their papers we have counts for

collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2023

GEAR: A GPU-Centric Experience Replay System for Large Reinforcement Learning Models

Hanjing Wang, Man-Kit Sit, Congjie He +5

This paper introduces a distributed, GPU-centric experience replay system, GEAR, designed to perform scalable reinforcement learning (RL) with large sequence models (such as transf…

cs.LG2023★ 4 cited

Large Sequence Models for Sequential Decision-Making: A Survey

Muning Wen, Runji Lin, Hanjing Wang +6

Transformer architectures have facilitated the development of large-scale and general-purpose sequence models for prediction tasks in natural language processing and computer visio…

cs.LG2023★ 12 cited

OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Jiaming Ji, Jiayi Zhou, Borong Zhang +7

AI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant s…

cs.LG2023★ 22 cited

Heterogeneous-Agent Reinforcement Learning

Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng +3

The necessity for cooperation among intelligent machines has popularised cooperative multi-agent reinforcement learning (MARL) in AI research. However, many research endeavours hea…

cs.LG2022★ 4 cited

ACE: Cooperative Multi-agent Q-learning with Bidirectional Action-Dependency

Chuming Li, Jie Liu, Yinmin Zhang +5

Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policie…

cs.LG2022★ 4 cited

Contextual Transformer for Offline Meta Reinforcement Learning

Runji Lin, Ye Li, Xidong Feng +6

The pretrain-finetuning paradigm in large-scale sequence models has made significant progress in natural language processing and computer vision tasks. However, such a paradigm is…