◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yu Zhang

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4
same name
  • Yu Zhang — 91 papers, h 38
  • Yu Zhang — 52 papers, h 51
  • Yu Zhang — 27 papers
  • Yu Zhang — 21 papers, h 18
  • Yu Zhang — 17 papers, h 32
  • Yu Zhang — 13 papers, h 32

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20192022
most citedCUP: A Conservative Update Policy Algorithm for Safe Reinforcement Learning

7 citations · 7 across the 2 of their papers we have counts for

collaborators

4 papers

cs.LG2022★ 7 cited

CUP: A Conservative Update Policy Algorithm for Safe Reinforcement Learning

Long Yang, Jiaming Ji, Juntao Dai +3

Safe reinforcement learning (RL) is still very challenging since it requires the agent to consider both return maximization and safe exploration. In this paper, we propose CUP, a C…

cs.LG2020

On Convergence of Gradient Expected Sarsa(λ)

Long Yang, Gang Zheng, Yu Zhang +3

We study the convergence of Expected Sarsa(λ) with linear function approximation. We show that applying the off-line estimate (multi-step bootstrapping) to $\mathtt{Expe…

cs.LG2019

Gradient Q(σ,λ): A Unified Algorithm with Function Approximation for Reinforcement Learning

Long Yang, Yu Zhang, Qian Zheng +2

Full-sampling (e.g., Q-learning) and pure-expectation (e.g., Expected Sarsa) algorithms are efficient and frequently used techniques in reinforcement learning. Q(σ,λ) is the firs…

cs.LG2019

Expected Sarsa(λ) with Control Variate for Variance Reduction

Long Yang, Yu Zhang, Jun Wen +3

Off-policy learning is powerful for reinforcement learning. However, the high variance of off-policy evaluation is a critical challenge, which causes off-policy learning falls into…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.