collaborators

5 papers

cs.LG2026

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation

Bo Xue, Zhi Hong, Jiayi Li +3

Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives. In this paper, we formulate…

cs.LG2026

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits

Bo Xue, Ji Cheng, Haodong Jing +2

This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued rew…

cs.LG2026

Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories

Dake Bu, Wei Huang, Andi Han +5

Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (…

cs.NE2025

Parametric Pareto Set Learning for Expensive Multi-Objective Optimization

Ji Cheng, Bo Xue, Qingfu Zhang

Parametric multi-objective optimization (PMO) addresses the challenge of solving an infinite family of multi-objective optimization problems, where optimal solutions must adapt to…

cs.LG2025

Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits

Bo Xue, Yuanyu Wan, Zhichao Lu +1

In multi-objective decision-making with hierarchical preferences, lexicographic bandits provide a natural framework for optimizing multiple objectives in a prioritized order. In th…