◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Pengfei Li

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4
same name
  • Pengfei Li — 13 papers, h 12
  • Pengfei Li — 11 papers, h 14
  • Pengfei Li — 8 papers, h 11
  • Pengfei Li — 6 papers, h 45
  • Pengfei Li — 3 papers
  • Pengfei Li — 3 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20182022
most citedCUP: A Conservative Update Policy Algorithm for Safe Reinforcement Learning

7 citations · 7 across the 2 of their papers we have counts for

collaborators

4 papers

cs.LG2022★ 7 cited

CUP: A Conservative Update Policy Algorithm for Safe Reinforcement Learning

Long Yang, Jiaming Ji, Juntao Dai +3

Safe reinforcement learning (RL) is still very challenging since it requires the agent to consider both return maximization and safe exploration. In this paper, we propose CUP, a C…

cs.LG2020

On Convergence of Gradient Expected Sarsa(λ)

Long Yang, Gang Zheng, Yu Zhang +3

We study the convergence of Expected Sarsa(λ) with linear function approximation. We show that applying the off-line estimate (multi-step bootstrapping) to $\mathtt{Expe…

cs.LG2019

Gradient Q(σ,λ): A Unified Algorithm with Function Approximation for Reinforcement Learning

Long Yang, Yu Zhang, Qian Zheng +2

Full-sampling (e.g., Q-learning) and pure-expectation (e.g., Expected Sarsa) algorithms are efficient and frequently used techniques in reinforcement learning. Q(σ,λ) is the firs…

cs.LG2018

Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network

Wenjia Meng, Qian Zheng, Long Yang +2

The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. DQN brings advances to complex sequential decision problems,…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.