◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Xiang Cheng

5 papers hereh-index 3192 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3
  • last author1

Across the 4 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.AI1
same name
  • Xiang Cheng — 17 papers, h 8
  • Xiang Cheng — 13 papers, h 6
  • Xiang Cheng — 8 papers, h 2
  • Xiang Cheng — 5 papers, h 3
  • Xiang Cheng — 5 papers, h 4
  • Xiang Cheng — 5 papers, h 9

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.LG2026

VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training

Guobin Shen, Chenxiao Zhao, Xiang Cheng +2

Off-policy updates are inevitable in reinforcement learning (RL) for large language models (LLMs) due to rollout staleness from asynchronous training and mismatches between trainin…

cs.LG2026

GLM-5: from Vibe Coding to Agentic Engineering

GLM-5-Team, :, Aohan Zeng +184

We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (AR…

cs.LG2026

LRT-Diffusion: Calibrated Risk-Aware Guidance for Diffusion Policies

Ximan Sun, Xiang Cheng

Diffusion policies are competitive for offline reinforcement learning (RL) but are typically guided at sampling time by heuristics that lack a statistical notion of risk. We introd…

cs.AI2025

Towards Agentic Self-Learning LLMs in Search Environment

Wangtao Sun, Xiang Cheng, Jialin Fan +5

We study whether self-learning can scale LLM-based agents without relying on human-curated datasets or predefined rule-based rewards. Through controlled experiments in a search-age…

cs.LG2025

Probabilistic Uncertain Reward Model

Wangtao Sun, Xiang Cheng, Xing Yu +5

Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.