activity
20242026
collaborators

5 papers

cs.LG2026

On the optimization dynamics of RLVR: Gradient gap and step size thresholds

Joe Suk, Yaqi Duan

Reinforcement Learning with Verifiable Rewards (RLVR), which uses simple binary feedback to post-train large language models, has found significant empirical success. However, a pr…

cs.AI2025

Ask, Clarify, Optimize: Human-LLM Agent Collaboration for Smarter Inventory Control

Yaqi Duan, Yichun Hu, Jiashuo Jiang

Inventory management remains a challenge for many small and medium-sized businesses that lack the expertise to deploy advanced optimization methods. This paper investigates whether…

cs.LG2025

Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting

Yunzhen Feng, Parag Jain, Anthony Hartshorn +2

Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for improving large language models (LLMs) on reasoning tasks, with Group Relative Policy Optimiz…

cs.LG2025

PILAF: Optimal Human Preference Sampling for Reward Modeling

Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng +2

As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerge…

stat.ML2024

Localized exploration in contextual dynamic pricing achieves dimension-free regret

Jinhang Chai, Yaqi Duan, Jianqing Fan +1

We study the problem of contextual dynamic pricing with a linear demand model. We propose a novel localized exploration-then-commit (LetC) algorithm which starts with a pure explor…