◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Shaohang Wei

Peking University

13 papers hereh-index 567 citations18 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author11

Across the 13 of 13 papers where every author was matched, so the position is known.

fields
  • cs.CL7
  • cs.LG4
  • cs.AI2
affiliations
  • Peking University
HomepageORCID 0009-0000-9931-7160

identity via Semantic Scholar / OpenAlex

most citedProbability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

Verifier-Induced Support Reshaping in On-Policy Optimization

Shaohang Wei, Zikun Su, Feifan Song +4

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sa…

cs.LG2026

Experience Augmented Policy Optimization for LLM Reasoning

Jinda Lu, Kexin Huang, Junkang Wu +7

Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…

cs.LG2026★ 1 cited

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

Wenhao Yu, Shaohang Wei, Jiahong Liu +5

Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dimensional: the ground-truth probability…

cs.LG2026

One-Way Policy Optimization for Self-Evolving LLMs

Shuo Yang, Jinda Lu, Kexin Huang +6

Reinforcement Learning with Verifiable Rewards (RLVR) has become a promising paradigm for scaling reasoning capabilities of Large Language Models (LLMs). However, the sparsity of b…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.