6 papers
Optimal and Order-optimal Gated Priority-based Greedy Policies for Two-layer Multi-item Order Fulfillment
Xi Chen, Yuze Chen, Ziyi Chen +1
We study how an e-commerce firm should make real-time fulfillment decisions in a two-layer distribution network when multi-item customer orders arrive sequentially and future deman…
Skill Weaving: Efficient LLM Improvement via Modular Skillpacks
Zhuo Li, Guodong Du, Zesheng Shi +5
Large language models increasingly require specialization across diverse domains, yet existing approaches struggle to balance multi-domain capacities with strict memory and inferen…
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
Wenjie Tang, Minne Li, Sijie Huang +2
Reinforcement learning from verifiable rewards (RLVR) is a promising paradigm for improving large language model (LLM) agents on long-horizon interactive tasks. However, in partial…
Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects
Xi Chen, Shibo Dai, Jiameng Lyu +1
We study the dynamic joint assortment selection and positioning problem, where the attraction of each product depends on both its intrinsic appeal and its display position under a…
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
Xi Chen, Yuze Chen, Yuan Zhou
We study a class of multi-period online decision-making problems with sequence-based predictions, which may be generated by machine learning models but whose accuracy is not guaran…
A Minimax-MDP Framework with Future-imposed Conditions for Learning-augmented Problems
Xin Chen, Yuze Chen, Yuan Zhou
We study a class of sequential decision-making problems with augmented predictions, potentially provided by a machine learning algorithm. In this setting, the decision-maker receiv…