◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yi-Ning Li

8 papers hereh-index 433 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author5

Across the 8 of 8 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.AI1
  • cs.CL1
  • cs.CV1
  • cs.RO1

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedHow to Find the Exact Pareto Front for Multi-Objective MDPs?

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang +5

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RL…

cs.LG2026

Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual

Yining Li, Peizhong Ju, Ness Shroff

Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constrai…

cs.LG2024

How to Find the Exact Pareto Front for Multi-Objective MDPs?

Yining Li, Peizhong Ju, Ness B. Shroff

Multi-Objective Markov Decision Processes (MO-MDPs) are receiving increasing attention, as real-world decision-making problems often involve conflicting objectives that cannot be a…

cs.LG2023

Achieving Sample and Computational Efficient Reinforcement Learning by Action Space Reduction via Grouping

Yining Li, Peizhong Ju, Ness Shroff

Reinforcement learning often needs to deal with the exponential growth of states and actions when exploring optimal control in high-dimensional spaces (often known as the curse of…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.