◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Hao Yang

6 papers hereh-index 212 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author3

Across the 6 of 6 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL2
same name
  • Hao Yang — 13 papers, h 6
  • Hao Yang — 10 papers, h 6
  • Hao Yang — 8 papers, h 8
  • Hao Yang — 7 papers, h 8
  • Hao Yang — 6 papers, h 3
  • Hao Yang — 6 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

works on
large language model evaluation 1preference learning 1query-only supervision 1rubric generation 1synthetic pairwise data 1

From the 1 of 6 linked papers with an AI index.

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

Uncertainty-Aware Reward Modeling for Stable RLHF

Licheng Pan, Haocheng Yang, Haoxuan Li +7

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…

cs.LG2026

Looped World Models

Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28

Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding error…

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.LG2025

Estimating the Effects of Sample Training Orders for Large Language Models without Retraining

Hao Yang, Haoxuan Li, Mengyue Yang +2

The order of training samples plays a crucial role in large language models (LLMs), significantly impacting both their external performance and internal learning dynamics. Traditio…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.