activity
20172025
most citedNeural Certificates for Safe Control Policies

43 citations · 289 across the 32 of their papers we have counts for

collaborators

57 papers

cs.LG2025

The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability

Jiachen Hu, Rui Ai, Han Zhong +4

Information asymmetry is a pervasive feature of multi-agent systems, especially evident in economics and social sciences. In these settings, agents tailor their actions based on pr…

cs.GT2024

An Instrumental Value for Data Production and its Application to Data Pricing

Rui Ai, Boxiang Lyu, Zhaoran Wang +2

How much value does a dataset or a data production process have to an agent who wishes to use the data to assist decision-making? This is a fundamental question towards understandi…

stat.ML2023

Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks

Jianqing Fan, Zhaoran Wang, Zhuoran Yang +1

We study high-dimensional multi-armed contextual bandits with batched feedback where the steps of online interactions are divided into batches. In specific, each batch coll…

cs.LG20221 cited

Relational Reasoning via Set Transformers: Provable Efficiency and Applications to MARL

Fengzhuo Zhang, Boyi Liu, Kaixin Wang +3

The cooperative Multi-A gent R einforcement Learning (MARL) with permutation invariant agents framework has achieved tremendous empirical successes in real-world applications. Unfo…

cs.LG20226 cited

Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Zuyue Fu, Zhengling Qi, Zhaoran Wang +3

We study the offline reinforcement learning (RL) in the face of unmeasured confounders. Due to the lack of online interaction with the environment, offline RL is facing the followi…

cs.LG20222 cited

Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation

Xiaoyu Chen, Han Zhong, Zhuoran Yang +2

We study human-in-the-loop reinforcement learning (RL) with trajectory preferences, where instead of receiving a numeric reward at each step, the agent only receives preferences ov…