activity
20222024
most citedHGAttack: Transferable Heterogeneous Graph Adversarial Attack

2 citations · 4 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2024

Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models

Fuxiang Zhang, Junyou Li, Yi-Chen Li +3

Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge…

cs.LG20242 cited

HGAttack: Transferable Heterogeneous Graph Adversarial Attack

He Zhao, Zhiwei Zeng, Yongwei Wang +2

Heterogeneous Graph Neural Networks (HGNNs) are increasingly recognized for their performance in areas like the web and e-commerce, where resilience against adversarial attacks is…

cs.LG2023

Master-slave Deep Architecture for Top-K Multi-armed Bandits with Non-linear Bandit Feedback and Diversity Constraints

Hanchi Huang, Li Shen, Deheng Ye +1

We propose a novel master-slave architecture to solve the top- combinatorial multi-armed bandits problem with non-linear bandit feedback and diversity constraints, which, to the…

cs.LG20232 cited

Future-conditioned Unsupervised Pretraining for Decision Transformer

Zhihui Xie, Zichuan Lin, Deheng Ye +3

Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…

cs.LG2023

Revisiting Estimation Bias in Policy Gradients for Deep Reinforcement Learning

Haoxuan Pan, Deheng Ye, Xiaoming Duan +4

We revisit the estimation bias in policy gradients for the discounted episodic Markov decision process (MDP) from Deep Reinforcement Learning (DRL) perspective. The objective is fo…

cs.LG2023

Sample Dropout: A Simple yet Effective Variance Reduction Technique in Deep Policy Optimization

Zichuan Lin, Xiapeng Wu, Mingfei Sun +4

Recent success in Deep Reinforcement Learning (DRL) methods has shown that policy optimization with respect to an off-policy distribution via importance sampling is effective for s…