◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Shuo Yang

10 papers hereh-index 450 citations12 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author8

Across the 10 of 10 papers where every author was matched, so the position is known.

fields
  • cs.CL4
  • cs.LG4
  • cs.AI1
  • cs.CV1
same name
  • Shuo Yang — 27 papers, h 11
  • Shuo Yang — 18 papers, h 14
  • Shuo Yang — 11 papers, h 7
  • Shuo Yang — 10 papers, h 5
  • Shuo Yang — 9 papers, h 2
  • Shuo Yang — 9 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

works on
agentic reinforcement learning 1on-policy distillation 1sample efficiency 1skill extraction 1text-based agents 1

From the 1 of 10 linked papers with an AI index.

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Kexin Huang, Junkang Wu, Jinda Lu +7

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this wor…

cs.LG2026

Experience Augmented Policy Optimization for LLM Reasoning

Jinda Lu, Kexin Huang, Junkang Wu +7

Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…

cs.LG2026

Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning

Jinyang Wu, Shuo Yang, Changpeng Yang +4

Reinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challenging due to the scarcity of high-qu…

cs.LG2026

FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization

Chiyu Ma, Shuo Yang, Kexin Huang +7

We present Future-KL Influenced Policy Optimization (FIPO), a reinforcement learning algorithm designed to overcome reasoning bottlenecks in large language models. While GRPO style…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.