activity
20242026
most citedWhat Matters in Data for DPO?

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

Hanzhao Wang, Jingxuan Wu, Yumeng Li +2

The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data center…

cs.LG2026

Learning Shortest Paths When Data is Scarce

Dmytro Matsypura, Yu Pan, Hanzhao Wang

Digital twins and other simulators are increasingly used to support routing decisions in large-scale networks. However, simulator outputs often exhibit systematic bias, while groun…

cs.LG2025

What Matters in Data for DPO?

Yu Pan, Zhongze Cai, Guanting Chen +2

Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…

cs.LG2024

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

Shang Liu, Yu Pan, Guanting Chen +1

Learning a reward model (RM) from human preferences has been an important component in aligning large language models (LLMs). The canonical setup of learning RMs from pairwise pref…

cs.LG2024

Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making

Hanzhao Wang, Yu Pan, Fupeng Sun +4

In this paper, we consider the supervised pre-trained transformer for a class of sequential decision-making problems. The class of considered problems is a subset of the general fo…

cs.LG2024

Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach

Linyu Liu, Yu Pan, Xiaocheng Li +1

In this paper, we study the problem of uncertainty estimation and calibration for LLMs. We begin by formulating the uncertainty estimation problem, a relevant yet underexplored are…