5 papers · 1 filter
Learning Shortest Paths When Data is Scarce
Dmytro Matsypura, Yu Pan, Hanzhao Wang
Digital twins and other simulators are increasingly used to support routing decisions in large-scale networks. However, simulator outputs often exhibit systematic bias, while groun…
What Matters in Data for DPO?
Yu Pan, Zhongze Cai, Guanting Chen +2
Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Shang Liu, Yu Pan, Guanting Chen +1
Learning a reward model (RM) from human preferences has been an important component in aligning large language models (LLMs). The canonical setup of learning RMs from pairwise pref…
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
Linyu Liu, Yu Pan, Xiaocheng Li +1
In this paper, we study the problem of uncertainty estimation and calibration for LLMs. We begin by formulating the uncertainty estimation problem, a relevant yet underexplored are…
Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making
Hanzhao Wang, Yu Pan, Fupeng Sun +4
In this paper, we consider the supervised pre-trained transformer for a class of sequential decision-making problems. The class of considered problems is a subset of the general fo…