15 papers
PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents
Zhenbang Du, Jun Luo, Zhiwei Zheng +8
Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning…
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
Yuchen Liang, Ness Shroff, Yingbin Liang
Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, but, especially for uniform-rate models, they often require many steps to g…
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
Ming Shi, Yingbin Liang, Ness B. Shroff +1
Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference la…
Learnable Chernoff Baselines for Inference-Time Alignment
Sunil Madhow, Yuchen Liang, Ness Shroff +2
We study inference-time reward-guided alignment for generative models. Existing methods often rely on either architecture-specific adaptations or computationally costly inference p…
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
Ming Shi, Yingbin Liang, Ness B. Shroff
Partially observable Markov decision processes (POMDPs) are a general framework for sequential decision-making under latent state uncertainty, yet learning in POMDPs is intractable…
Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
Yuchen Liang, Yingbin Liang, Lifeng Lai +1
Discrete diffusion models have recently gained significant prominence in applications involving natural language and graph data. A key factor influencing their effectiveness is the…