9 papers
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…
Adaptive Multi-Round Allocation with Stochastic Arrivals
Yuqi Pan, Davin Choo, Haichuan Wang +3
We study a sequential resource allocation problem motivated by adaptive network recruitment, in which a limited budget of identical resources must be allocated over multiple rounds…
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
Lingkai Kong, Haichuan Wang, Tonghan Wang +2
Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamic…
The Publication Choice Problem
Haichuan Wang, Yifan Wu, Haifeng Xu
Researchers strategically choose where to submit their work in order to maximize its impact, and these publication decisions in turn determine venues' impact factors. To analyze ho…
Robust Optimization with Diffusion Models for Green Security
Lingkai Kong, Haichuan Wang, Yuqi Pan +6
In green security, defenders must forecast adversarial behavior, such as poaching, illegal logging, and illegal fishing, to plan effective patrols. These behavior are often highly…
Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting
Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar +6
Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union sta…