4 papers
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…
Generative Frontier Planning for Adaptive Peer-Referral Recruitment under Covariate-Dependent Arrivals
Lingkai Kong, Hezi Jiang, Andrew Ma +3
Peer-referral recruitment systems such as respondent-driven sampling are critical for studying and intervening on hidden populations affected by infectious diseases. To accelerate…
Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
Lingkai Kong, Anagha Satish, Hezi Jiang +6
Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraint…
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
Sirui Xu, Dongting Li, Yucheng Zhang +9
While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to…