6 papers
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification
Haoyang Hong, Zichen Wang, Quanquan Gu +1
We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on re…
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
Jose Efraim Aguilar Escamilla, Haoyang Hong, Jiawei Li +4
We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards within constrained budgets to force the target RL agent to adopt a policy t…
Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies
Muyun Lu, Haoyang Hong, Huazheng Wang +1
As an extension of the classical multi-armed bandit problem, multi-fidelity multi-armed bandits (MF-MAB) enable individual arms to be evaluated using diverse feedback sources that…
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
Haoyang Hong, Jiajun Yin, Yuan Wang +14
Multi-agent systems perform well on general reasoning tasks. However, the lack of training in specialized areas hinders their accuracy. Current training methods train a unified lar…
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
Zichen Wang, Haoyang Hong, Chuanhao Li +3
In multi-armed bandits with network interference (MABNI), the action taken by one node can influence the rewards of others, creating complex interdependence. While existing researc…
Do regularization methods for shortcut mitigation work as intended?
Haoyang Hong, Ioanna Papanikolaou, Sonali Parbhoo
Mitigating shortcuts, where models exploit spurious correlations in training data, remains a significant challenge for improving generalization. Regularization methods have been pr…