3 papers
cs.LG2026
Learning Shortest Paths When Data is Scarce
Dmytro Matsypura, Yu Pan, Hanzhao Wang
Digital twins and other simulators are increasingly used to support routing decisions in large-scale networks. However, simulator outputs often exhibit systematic bias, while groun…
cs.LG2025
What Matters in Data for DPO?
Yu Pan, Zhongze Cai, Guanting Chen +2
Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…
cs.SE2025
Online-Optimized RAG for Tool Use and Function Calling
Yu Pan, Xiaocheng Li, Hanzhao Wang
In many applications, retrieval-augmented generation (RAG) drives tool use and function calling by embedding the (user) queries and matching them to pre-specified tool/function des…