◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yang Xu

14 papers hereh-index 457 citations16 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author8
  • middle author5

Across the 13 of 14 papers where every author was matched, so the position is known.

fields
  • cs.LG6
  • stat.ML4
  • cs.AI2
  • cs.RO1
  • quant-ph1
same name
  • Yang Xu — 10 papers, h 2
  • Yang Xu — 7 papers, h 1
  • Yang Xu — 7 papers, h 7
  • Yang Xu — 7 papers, h 25
  • Yang Xu — 6 papers, h 1
  • Yang Xu — 5 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

works on
actor-critic 1average-reward MDP 1distributional robustness 1q-learning 1robust reinforcement learning 1

From the 1 of 14 linked papers with an AI index.

collaborators
Showing stat.MLShow all

4 papers · 1 filter

stat.ML2026

Core-Halo Decomposition: Decentralizing Large-Scale Fixed-Point Problems

Haixiang, Yang Xu, Jiefu Zhang +4

We study solving large-scale fixed-point equation \(x^\star=\bar F(x^\star)\) with decomposition. Standard strict decomposition assigns each agent a disjoint block and evaluates up…

stat.ML2026

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

Yang Xu, Jiefu Zhang, Haixiang Sun +3

Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…

stat.ML2026

Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients

Yang Xu, Vaneet Aggarwal

We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnosti…

stat.ML2025

Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning

Yang Xu, Washim Uddin Mondal, Vaneet Aggarwal

We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asympto…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.