◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Kai Yang

10 papers hereh-index 468 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author8
  • last author1

Across the 10 of 10 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL3
  • cs.AI1
  • cs.CV1
  • cs.DB1
same name
  • Kai Yang — 11 papers, h 5
  • Kai Yang — 6 papers, h 2
  • Kai Yang — 6 papers, h 4
  • Kai Yang — 6 papers, h 2
  • Kai Yang — 5 papers, h 2
  • Kai Yang — 5 papers, h 11

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data

Hsiu-Yuan Huang, Weijie Liu, Chenming Tang +5

The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among existing datasets. To bridge th…

cs.LG2026

Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Wenkai Yang, Weijie Liu, Ruobing Xie +3

On-policy distillation (OPD), which aligns the student with the teacher's logit distribution on student-generated trajectories, has demonstrated strong empirical gains in improving…

cs.LG2026

EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control

Kai Yang, Xin Xu, Yangkun Chen +5

Long-term training of large language models (LLMs) requires maintaining stable exploration to prevent the model from collapsing into sub-optimal behaviors. Entropy is crucial in th…

cs.LG2025

Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners

Xin Xu, Cliveb AI, Kai Yang +4

Reinforcement Learning with Verifiable Reward (RLVR) effectively solves complex tasks but demands extremely long context lengths during training, leading to substantial computation…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.