5 papers
Learning Shortest Paths When Data is Scarce
Dmytro Matsypura, Yu Pan, Hanzhao Wang
Digital twins and other simulators are increasingly used to support routing decisions in large-scale networks. However, simulator outputs often exhibit systematic bias, while groun…
Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated Learning
Zhuang Qi, Pan Yu, Lei Meng +4
Federated Prompt Learning (FPL) enables communication-efficient adaptation by tuning lightweight prompts on top of frozen pre-trained models. Existing FPL methods typically rely on…
Online-Optimized RAG for Tool Use and Function Calling
Yu Pan, Xiaocheng Li, Hanzhao Wang
In many applications, retrieval-augmented generation (RAG) drives tool use and function calling by embedding the (user) queries and matching them to pre-specified tool/function des…
What Matters in Data for DPO?
Yu Pan, Zhongze Cai, Guanting Chen +2
Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Shang Liu, Yu Pan, Guanting Chen +1
Learning a reward model (RM) from human preferences has been an important component in aligning large language models (LLMs). The canonical setup of learning RMs from pairwise pref…