activity
20242026
most citedRethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Xinyu Tang, Qianggang Cao, Yurou Liu +13

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoni…

cs.CL2026

GraphPO: Graph-based Policy Optimization for Reasoning Models

Yuliang Zhan, Xinyu Tang, Jian Li +7

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…

cs.CL20251 cited

Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

Xinyu Tang, Yuliang Zhan, Zhixun Li +5

Large reasoning models (LRMs) are typically trained using reinforcement learning with verifiable reward (RLVR) to enhance their reasoning abilities. In this paradigm, policies are…

cs.CL2025

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

Yuliang Zhan, Xinyu Tang, Han Wan +3

Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision-Language Models (VLMs) still struggle with multi-…

cs.CL2025

Enhancing Cross-task Transfer of Large Language Models via Activation Steering

Xinyu Tang, Zhihao Lv, Xiaoxue Cheng +5

Large language models (LLMs) have shown impressive abilities in leveraging pretrained knowledge through prompting, but they often struggle with unseen tasks, particularly in data-s…

cs.CL2025

Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering

Xinyu Tang, Xiaolei Wang, Zhihao Lv +5

Recent advancements in long chain-of-thoughts(long CoTs) have significantly improved the reasoning capabilities of large language models(LLMs). Existing work finds that the capabil…