activity
20182023
most citedNeural Certificates for Safe Control Policies

43 citations · 247 across the 28 of their papers we have counts for

collaborators

37 papers

q-fin.TR20228 cited

FinRL-Meta: Market Environments and Benchmarks for Data-Driven Financial Reinforcement Learning

Xiao-Yang Liu, Ziyi Xia, Jingyang Rui +6

Finance is a particularly difficult playground for deep reinforcement learning. However, establishing high-quality market environments and benchmarks for financial reinforcement le…

cs.LG20221 cited

Relational Reasoning via Set Transformers: Provable Efficiency and Applications to MARL

Fengzhuo Zhang, Boyi Liu, Kaixin Wang +3

The cooperative Multi-A gent R einforcement Learning (MARL) with permutation invariant agents framework has achieved tremendous empirical successes in real-world applications. Unfo…

cs.LG20226 cited

Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Zuyue Fu, Zhengling Qi, Zhaoran Wang +3

We study the offline reinforcement learning (RL) in the face of unmeasured confounders. Due to the lack of online interaction with the environment, offline RL is facing the followi…

cs.LG20222 cited

Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Xiaoyu Chen, Han Zhong, Zhuoran Yang +2

We study human-in-the-loop reinforcement learning (RL) with trajectory preferences, where instead of receiving a numeric reward at each step, the agent only receives preferences ov…

cs.LG2022

Learn to Match with No Regret: Reinforcement Learning in Markov Matching Markets

Yifei Min, Tianhao Wang, Ruitu Xu +3

We study a Markov matching market involving a planner and a set of strategic agents on the two sides of the market. At each step, the agents are presented with a dynamical context,…

cs.LG202224 cited

Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Chenjia Bai, Lingxiao Wang, Zhuoran Yang +4

Offline Reinforcement Learning (RL) aims to learn policies from previously collected datasets without exploring the environment. Directly applying off-policy algorithms to offline…