activity
20242026
collaborators

38 papers

cs.SI2026

Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples

Akseli Kangaslahti, Davin Choo, Lingkai Kong +3

HIV is a retrovirus that attacks the human immune system and can lead to death without proper treatment. In collaboration with the WHO and the University of Witwatersrand, we study…

cs.LG2026

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

Haichuan Wang, Tao Lin, Lingkai Kong +3

Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…

cs.LG2026

Generative Frontier Planning for Adaptive Peer-Referral Recruitment under Covariate-Dependent Arrivals

Lingkai Kong, Hezi Jiang, Andrew Ma +3

Peer-referral recruitment systems such as respondent-driven sampling are critical for studying and intervening on hidden populations affected by infectious diseases. To accelerate…

cs.LG2026

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

Lingkai Kong, Anagha Satish, Hezi Jiang +6

Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraint…

cs.AI2026

Generating Robust Portfolios of Optimization Models using Large Language Models

Eleni Straitouri, Cheol Woo Kim, Milind Tambe

Mathematical optimization is a powerful tool for structured decision-making across domains such as resource allocation and planning. Formulating optimization models faithful to rea…

cs.LG2026

Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning

Shresth Verma, Mauricio Tec, Cheol Woo Kim +2

While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a scalable approach, yet its perfor…