activity
20182023
most citedSMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

103 citations · 467 across the 30 of their papers we have counts for

collaborators

46 papers

cs.LG202113 cited

Dynamic Bottleneck for Robust Self-Supervised Exploration

Chenjia Bai, Lingxiao Wang, Lei Han +4

Exploration methods based on pseudo-count of transitions or curiosity of dynamics have achieved promising results in solving reinforcement learning with sparse rewards. However, su…

cs.LG202124 cited

Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual Learning

Danruo Deng, Guangyong Chen, Jianye Hao +2

The backpropagation networks are notably susceptible to catastrophic forgetting, where networks tend to forget previously learned skills upon learning new ones. To address such the…

cs.IR20212 cited

CMML: Contextual Modulation Meta Learning for Cold-Start Recommendation

Xidong Feng, Chen Chen, Dong Li +3

Practical recommender systems experience a cold-start problem when observed user-item interactions in the history are insufficient. Meta learning, especially gradient based one, ca…

cs.AI202110 cited

Cooperative Multi-Agent Transfer Learning with Level-Adaptive Credit Assignment

Tianze Zhou, Fubiao Zhang, Kun Shao +10

Extending transfer learning to cooperative multi-agent reinforcement learning (MARL) has recently received much attention. In contrast to the single-agent setting, the coordination…

cs.LG20211 cited

Contrastive ACE: Domain Generalization Through Alignment of Causal Mechanisms

Yunqi Wang, Furui Liu, Zhitang Chen +4

Domain generalization aims to learn knowledge invariant across different distributions while semantically meaningful for downstream tasks from multiple source domains, to improve t…

cs.LG20217 cited

Principled Exploration via Optimistic Bootstrapping and Backward Induction

Chenjia Bai, Lingxiao Wang, Lei Han +4

One principled approach for provably efficient exploration is incorporating the upper confidence bound (UCB) into the value function as a bonus. However, UCB is specified to deal w…