collaborators

11 papers

cs.LG2026

Optimal and Order-optimal Gated Priority-based Greedy Policies for Two-layer Multi-item Order Fulfillment

Xi Chen, Yuze Chen, Ziyi Chen +1

We study how an e-commerce firm should make real-time fulfillment decisions in a two-layer distribution network when multi-item customer orders arrive sequentially and future deman…

cs.LG2026

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability

Yushen Liu, Yin-Jen Chen, Ziyi Chen +4

Many safety-critical control problems are modeled as risk-sensitive partially observable Markov decision processes, where the controller must make decisions from incomplete observa…

cs.LG2026

Distributionally Robust Multi-Objective Optimization

Yufeng Yang, Fangning Zhuo, Ziyi Chen +2

Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, the existing MOO formulations do not expli…

cs.LG2026

Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization

Anirudh Satheesh, Ziyi Chen, Furong Huang +1

We study robust Markov decision processes (RMDPs) with general policy parameterization under s-rectangular and non-rectangular uncertainty sets. Prior work is largely limited to ta…

cs.LG2025

Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment

Ziyi Chen, Junyi Li, Peiran Yu +1

Reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) are important techniques to align large language models (LLM) with human preference. Howe…

math.OC2025

Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization

Ziyi Chen, Peiran Yu, Heng Huang

This work aims to solve a stochastic nonconvex nonsmooth composite optimization problem. Previous works on composite optimization problem requires the major part to satisfy Lipschi…