9 citations · 15 across the 7 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2023
Query-Policy Misalignment in Preference-Based Reinforcement Learning
Xiao Hu, Jianxiong Li, Xianyuan Zhan +2
Preference-based reinforcement learning (PbRL) provides a natural way to align RL agents' behavior with human desired outcomes, but is often restrained by costly human feedback. To…
cs.LG2023★ 1 cited
Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Jianxiong Li, Xiao Hu, Haoran Xu +4
Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to des…
cs.LG2021
An Actor-Critic Method for Simulation-Based Optimization
Kuo Li, Qing-Shan Jia, Jiaqi Yan
We focus on a simulation-based optimization problem of choosing the best design from the feasible space. Although the simulation model can be queried with finite samples, its inter…