4 papers
LLM Active Alignment: A Nash Equilibrium Perspective
Tonghan Wang, Yuqi Pan, Xinyi Yang +3
We develop a game-theoretic framework for predicting and steering the behavior of populations of large language models (LLMs) through Nash equilibrium (NE) analysis. To avoid the i…
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
Lingkai Kong, Haichuan Wang, Tonghan Wang +2
Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamic…
Adaptive Frontier Exploration on Graphs with Applications to Network-Based Disease Testing
Davin Choo, Yuqi Pan, Tonghan Wang +3
We study a sequential decision-making problem on a -node graph where each node has an unknown label from a finite set , drawn from a joint distribution $…
Robust Optimization with Diffusion Models for Green Security
Lingkai Kong, Haichuan Wang, Yuqi Pan +6
In green security, defenders must forecast adversarial behavior, such as poaching, illegal logging, and illegal fishing, to plan effective patrols. These behavior are often highly…